Ornith-1.0-35B

DeepReinforce AI

Parameters

35.0B total / 3.0B active

MoE: total / active

Architecture

MoE decoder-only (qwen35moe), post-trained on Qwen 3.5; reasoning model with tool-calling. 35B total / ~3B active (256 experts, 8 routed + 1 shared).

Released

21.06.2026

License

MIT License

Open Weights Commercial Use Multimodal BF16 Ornith en

Input Modalities

text

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

262,144 tokens

Openness Index Score 100.0/100

About

Ornith-1.0-35B (GGUF builds at deepreinforce-ai/Ornith-1.0-35B; card at ornith-ai/Ornith-1.0-35B) is the lightweight 35B-MoE member of Ornith 1.0, the self-improving family of open-source models for agentic coding from deepreinforce-ai, designed for efficient single-GPU deployment. Ornith 1.0 ships in 9B-Dense, 31B-Dense, 35B-MoE and 397B-MoE variants post-trained on top of Gemma 4 and Qwen 3.5, achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.

Its signature is the self-improving training framework: RL teaches the model to generate not only solution rollouts but also the scaffolds that drive them - jointly optimizing scaffold-rollout co-optimization so the model discovers better search trajectories and higher-quality solutions. This record's lineage: post-trained with RL (self-scaffolding) on top of the Qwen3.5 35B base. The underlying model is a Qwen3.5-style sparse Mixture-of-Experts with 35B total and 3B active parameters: 40 layers mixing 30 linear-attention layers (Gated DeltaNet-style) with 10 full-attention layers (interval 4), 256 routed experts with 8 active per token plus 1 shared expert, one Multi-Token Prediction layer, 248K vocabulary and a 262,144-token context, with an optional Qwen3.5 ViT for image input on the multimodal builds.

Serving is optimized for GGUF runtimes: llama.cpp, Ollama, LM Studio and vLLM/SGLang OpenAI-compatible endpoints, with agent frameworks (Hermes, OpenClaw, OpenHands, opencode) and Unsloth for local fine-tuning. Released June 21, 2026 under the MIT License.

Training Data Post-trained (RL with self-scaffolding) on top of Qwen 3.5 35B base. RL jointly optimizes scaffold and solution rollouts for agentic coding.

Benchmark Scores

Benchmark Score Date
Terminal-Bench 2.1 (Terminus-2)
coding_agent
63.72%
25.06.2026
Terminal-Bench 2.1 (Claude Code)
coding_agent
88.51%
25.06.2026
SWE-bench Verified
coding_agent
86.24%
25.06.2026
SWE-bench Pro
coding_agent
63.00%
25.06.2026
SWE-bench Multilingual
coding_agent
76.09%
25.06.2026
NL2Repo
coding_agent
37.38%
25.06.2026
Claw-Eval Avg
coding_agent
84.19%
25.06.2026
SWE Atlas - QnA
coding_agent
50.91%
25.06.2026
SWE Atlas - RF
coding_agent
45.60%
25.06.2026
SWE Atlas - TW
coding_agent
36.39%
25.06.2026

Model Tree and Spaces

Model tree for ornith-ai/Ornith-1.0-35B

Adapters

1 model

Finetunes

17 models

Merges

12 models

Quantizations

179 models

Spaces using ornith-ai/Ornith-1.0-35B 3

Collection including ornith-ai/Ornith-1.0-35B

[

Ornith-1.0

Collection

Ornith-1.0 is  a family of open-source LLMs specialized for agentic coding. • 8 items • Updated 23 days ago • 394

](https://huggingface.co/collections/ornith-ai/ornith-10)

Citation

Citation

If you find our work helpful, feel free to give us a cite.

@misc{ornith-35b,
    title = {{Ornith-1.0-35B}: Agentic Coding, Open to All},
    url = {https://deep-reinforce.com/ornith_1_0.html},
    author = {{DeepReinforce Team}},
    year = {2026}
}

Safetensors

Model size

665k params

Tensor type

BF16

·

Agentic Usage: agent frameworks (Hermes, OpenClaw, OpenHands, opencode, Unsloth)

Agentic Usage

Ornith-1.0-35B excels in tool-calling and agentic coding capabilities.

Agent Frameworks

Because Ornith-1.0-35B exposes an OpenAI-compatible endpoint with tool calling, it works out of the box with standard agent frameworks. Below is a minimal example that connects Ornith-1.0-35B to tools through an MCP server.

import os
from openai import OpenAI

client = OpenAI(
    base_url=os.getenv("OPENAI_BASE_URL", "http://localhost:8000/v1"),
    api_key=os.getenv("OPENAI_API_KEY", "EMPTY"),
)

tools = [
    {
        "type": "function",
        "function": {
            "name": "run_shell",
            "description": "Run a shell command and return its output.",
            "parameters": {
                "type": "object",
                "properties": {
                    "command": {"type": "string", "description": "The command to run"}
                },
                "required": ["command"],
            },
        },
    }
]

messages = [{"role": "user", "content": "List the Python files in the current directory."}]

response = client.chat.completions.create(
    model="deepreinforce-ai/Ornith-1.0-35B",
    messages=messages,
    tools=tools,
    temperature=0.6,
    top_p=0.95,
)
print(response.choices[0].message)

Examples of using Ornith with agent harness:

Hermes Agent

## Hermes talks to any OpenAI-compatible endpoint — point it at your Ornith server.
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export MODEL="deepreinforce-ai/Ornith-1.0-35B"

Atomic.chat/ Ollama / llama.cpp

## Both runtimes load a GGUF build of Ornith (publish one at deepreinforce-ai/Ornith-1.0-35B-GGUF).

## llama.cpp — serve an OpenAI-compatible API on port 8000.
llama-server -hf deepreinforce-ai/Ornith-1.0-35B-GGUF --port 8000 -c 262144

## Ollama — pull and chat with the same GGUF straight from Hugging Face.
ollama run hf.co/deepreinforce-ai/Ornith-1.0-35B-GGUF

OpenClaw

## OpenClaw talks to any OpenAI-compatible endpoint — point it at your Ornith server.
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export OPENAI_MODEL="deepreinforce-ai/Ornith-1.0-35B"

Unsloth Studio

pip install unsloth

## Load Ornith for fast local inference or fine-tuning (Python):
##   from unsloth import FastLanguageModel
##   model, tokenizer = FastLanguageModel.from_pretrained(
##       "deepreinforce-ai/Ornith-1.0-35B",
##       max_seq_length=262144,
##       load_in_4bit=True,
##   )

OpenHands

pip install openhands-ai

## OpenHands routes through LiteLLM; the "openai/" prefix selects the OpenAI-compatible path.
export LLM_MODEL="openai/deepreinforce-ai/Ornith-1.0-35B"
export LLM_BASE_URL="http://localhost:8000/v1"
export LLM_API_KEY="EMPTY"

## Launch the CLI (or run the official OpenHands Docker image with the same env vars).
openhands

Coding CLIs

Ornith-1.0-35B is optimized for terminal-based coding agents. Point any OpenAI-compatible coding CLI at your Ornith-1.0-35B endpoint (set OPENAI_BASE_URL and OPENAI_API_KEY) to understand large codebases, automate tedious work, and ship faster.

OpenCode

## Register your local Ornith endpoint as a provider in ~/.config/opencode/opencode.json:
##
## {
##   "$schema": "https://opencode.ai/config.json",
##   "provider": {
##     "ornith": {
##       "npm": "@ai-sdk/openai-compatible",
##       "name": "Ornith (local)",
##       "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "EMPTY" },
##       "models": { "deepreinforce-ai/Ornith-1.0-35B": { "name": "Ornith-1.0-35B" } }
##     }
##   }
## }

opencode

Citation

If you find our work helpful, feel free to give us a cite.

@misc{ornith-35b,
    title = {{Ornith-1.0-35B}: Agentic Coding, Open to All},
    url = {https://deep-reinforce.com/ornith_1_0.html},
    author = {{DeepReinforce Team}},
    year = {2026}
}

Safetensors

Model size

665k params

Tensor type

BF16

·

Serving Ornith-1.0-35B (reasoning content, tool calls)

Serving Ornith-1.0-35B

The two recipes below stand up an OpenAI-compatible server on a single 8×80GB GPU node (tensor-parallel 8). Adjust --tensor-parallel-size / --tp to the number of GPUs you have.

vLLM

vllm serve deepreinforce-ai/Ornith-1.0-35B \
    --served-model-name Ornith-1.0-35B \
    --tensor-parallel-size 8 \
    --host 0.0.0.0 --port 8000 \
    --max-model-len 262144 \
    --gpu-memory-utilization 0.90 \
    --enable-prefix-caching \
    --enable-auto-tool-choice --tool-call-parser qwen3_xml \
    --reasoning-parser qwen3 \
    --trust-remote-code

SGLang

python -m sglang.launch_server \
    --model-path deepreinforce-ai/Ornith-1.0-35B \
    --served-model-name Ornith-1.0-35B \
    --tp 8 \
    --host 0.0.0.0 --port 8000 \
    --context-length 262144 \
    --mem-fraction-static 0.85 \
    --tool-call-parser qwen3_coder \
    --reasoning-parser qwen3

Hugging Face Transformers

For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-35B requires transformers >= 5.8.1.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "deepreinforce-ai/Ornith-1.0-35B"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    dtype="auto",
    device_map="auto",
)

messages = [
    {"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

inputs = tokenizer(text, return_tensors="pt").to(model.device)
generated = model.generate(
    **inputs,
    max_new_tokens=512,
    do_sample=True,
    temperature=0.6,
    top_p=0.95,
    top_k=20,
)
output_ids = generated[0][inputs.input_ids.shape[1]:]

## The reply contains a <think> ... </think> reasoning block followed by the answer.
content = tokenizer.decode(output_ids, skip_special_tokens=True)
print(content)

To split the reasoning trace from the final answer, parse on the </think> marker:

text = tokenizer.decode(output_ids, skip_special_tokens=True)
if "</think>" in text:
    reasoning, answer = text.split("</think>", 1)
    reasoning = reasoning.replace("<think>", "").strip()
    answer = answer.strip()
else:
    reasoning, answer = "", text.strip()

Quickstart

Quickstart

📝 NOTE

Ornith-1.0-35B is a reasoning model: by default the assistant turn opens with a <think> … </think> block before the final answer. The serving recipes below enable a reasoning parser so the chain-of-thought is returned in a separate reasoning_content field, and a tool-call parser so the model's <tool_call> blocks are surfaced as OpenAI-style tool_calls.

Serving Ornith-1.0-35B requires recent runtimes:

  • Transformers ≥ 5.8.1
  • vLLM ≥ 0.19.1
  • SGLang ≥ 0.5.9

Benchmarks (vs Qwen3.5-35B, Qwen3.6-35B, Gemma4-31B, Qwen3.5-397B)

Benchmarks

Ornith-1.0-35B Qwen3.5-35B Qwen3.6-35B Gemma4-31B Qwen3.5-397B
Agentic Coding
Terminal-Bench 2.1 (Terminus-2) 64.2 41.4 52.5 42.1 53.5
Terminal-Bench 2.1 (Claude Code) 62.8 38.9 49.2 - 48.6
SWE-bench Verified 75.6 70 73.4 52 76.4
SWE-bench Pro 50.4 44.6 49.5 35.7 51.6
SWE-bench Multilingual 69.3 60.3 67.2 51.7 69.3
NL2Repo 34.6 20.5 29.4 15.5 36.8
Claw-eval Avg 69.8 65.4 68.7 48.5 70.7
SWE Atlas - QnA 37.1 13.2 15.5 - 20.4
SWE Atlas - RF 29.7 10.2 11.4 - 18.4
SWE Atlas - TW 27.8 9.8 13.3 - 18.5

* Terminal-Bench 2.1 (Terminus-2): We evaluate Terminal-Bench 2.1 using the Harbor/Terminus-2 framework with parser=json, temperature=1.0, top_p=1.0, and a 128K context window. Each run uses a 4-hour timeout with 32 CPU cores and 48GB RAM, and results are averaged over 5 runs. We adjust the Qwen chat template to ensure consistency between training and inference (https://huggingface.co/deepreinforce-ai/Ornith-1.0-397B/blob/main/chat_template.jinja), and modify Harbor to align with vLLM's reasoning_content key.
* Terminal-Bench 2.1 (Claude Code): We evaluate Terminal-Bench 2.1 using Claude Code 2.1.126 with parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072. Results are averaged over 5 runs. Again, Qwen chat template needs to be modified.
* SWE-Bench Verified, Pro and Multilingual: using OpenHands harness with temp=1.0, top_p=0.95, 256k context window.
* SWE Atlas QnA, RF, TW: using mini SWE agent harness with temp=1.0, top_p=0.95, 128K context window. Results are averaged over 5 runs.
* NL2Repo: with temperature=1.0, top_p=1.0, 400K context, 48K output and anti-hacking filters.
* ClawEval: An agentic code benchmark over real-user task distributions; temp=0.6 and 256K context.

Ornith-1.0-35B Introduction (self-improving agentic coding)

Ornith-1.0-35B

Aloha! 🌺 Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding.

Highlights:

  • State-of-the-Art Coding Agents: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
  • Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scallfold that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.
  • Licence: MIT licensed, globally accessible, and free from regional limitations.

Ornith 35B Benchmark Results

Ornith 1.0 35B

This model card documents Ornith-1.0-35B, the lightweight member of the Ornith family, designed for efficient single-GPU deployment.

Architecture

Decoder Block ×40 input Embedding vocab 248K · d 2048 Linear / Recurrent Hybrid 16:2 · dₕ 256 ×30 Full Attention Hybrid 16:2 · dₕ 256 ×10 MoE FFN 256 experts · top-8 · dᴻ 512 Final Norm LM Head vocab 248K output
Attention
Hybrid Attention (16:2)
MoE
256 experts · top-8 per token
Layers
40
Hidden size
2048
Context
262K tokens
Parameters
35000M
Active params
3000M

Source: Hugging Face config.json · Qwen3_5MoeForConditionalGeneration · exact layer pattern · model repo

Type: decoder-only transformer
Attention: GQA
Decoder: causal
MoE: yes (256 experts)
Routing: top-8 routed + 1 shared
Total parameters 35000M
Active parameters 3000M
Context length 262K
Experts 256
Routed experts 8
Experts per token 9
Shared experts 1

Training Pipeline

  1. 1
    rl

    RL self-scaffolding post-training on Qwen 3.5 35B

    Ornith-1.0 employs RL to jointly optimize scaffold and solution rollouts for agentic coding. Post-trained on top of Qwen 3.5 35B base.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Agentic coding RL rollouts (solution + scaffold) rl — —

Trend Analysis

24h Change

+0.9%

7d Change

+9.3%

Current

3,185

huggingface

downloads

+0.4%

huggingface

downloads_all_time

+1.2%

huggingface

likes

+0.0%

ollama

downloads

+0.1%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 3,239 pulls daily 01.09.2026
huggingface downloads_all_time 6,075,393 daily 01.09.2026
huggingface followers 3,185 daily 01.09.2026
huggingface likes 504 daily 01.09.2026
huggingface downloads 3,056,578 daily 01.09.2026
ollama downloads 3,236 pulls daily 31.08.2026
huggingface downloads_all_time 6,003,553 daily 31.08.2026
huggingface followers 3,158 daily 31.08.2026
huggingface likes 504 daily 31.08.2026
huggingface downloads 3,045,441 daily 31.08.2026
ollama downloads 3,231 pulls daily 30.08.2026
huggingface downloads_all_time 5,952,173 daily 30.08.2026
huggingface followers 3,129 daily 30.08.2026
huggingface likes 504 daily 30.08.2026
huggingface downloads 3,034,579 daily 30.08.2026
ollama downloads 3,227 pulls daily 29.08.2026
huggingface followers 3,096 daily 29.08.2026
huggingface likes 504 daily 29.08.2026
huggingface downloads 2,994,126 daily 29.08.2026
ollama downloads 3,218 pulls daily 28.08.2026
huggingface followers 3,061 daily 28.08.2026
huggingface likes 503 daily 28.08.2026
huggingface downloads 3,005,458 daily 28.08.2026
ollama downloads 3,207 pulls daily 27.08.2026
huggingface followers 3,020 daily 27.08.2026
huggingface likes 503 daily 27.08.2026
huggingface downloads 3,127,649 daily 27.08.2026
ollama downloads 3,197 pulls daily 26.08.2026
huggingface followers 2,976 daily 26.08.2026
huggingface likes 504 daily 26.08.2026
huggingface downloads 3,234,380 daily 26.08.2026
ollama downloads 3,186 pulls daily 25.08.2026
huggingface followers 2,913 daily 25.08.2026
huggingface likes 505 daily 25.08.2026
huggingface downloads 3,286,288 daily 25.08.2026
ollama downloads 3,181 pulls daily 24.08.2026
huggingface followers 2,860 daily 24.08.2026
huggingface likes 505 daily 24.08.2026
huggingface downloads 3,323,741 daily 24.08.2026
huggingface followers 2,803 daily 23.08.2026
huggingface likes 505 daily 23.08.2026
huggingface downloads 3,188,828 daily 23.08.2026
huggingface followers 2,741 daily 22.08.2026
huggingface likes 505 daily 22.08.2026
huggingface downloads 3,192,730 daily 22.08.2026
huggingface followers 2,673 daily 21.08.2026
huggingface likes 505 daily 21.08.2026
huggingface downloads 3,228,856 daily 21.08.2026
huggingface followers 2,560 daily 20.08.2026
huggingface likes 504 daily 20.08.2026

View full metric history →

Related Models