Parameters
35.0B total / 3.0B active
MoE: total / active
Architecture
Mixture-of-Experts (qwen3_5_moe)
Released
01.07.2026
License
MIT License
Input Modalities
Output Modalities
Context (native)
262,144 tokens
Context (extended)
262,144 tokens
About
Ornith-1.0-35B (ornith-ai/Ornith-1.0-35B) is the lightweight 35B MoE member of Ornith 1.0, deepreinforce-ai's self-improving family of open-source models for agentic coding, designed for efficient single-GPU deployment. Ornith 1.0 ships in 9B-Dense, 31B-Dense, 35B-MoE and 397B-MoE variants post-trained on top of Gemma 4 and Qwen 3.5, achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
Its signature is the self-improving training framework: reinforcement learning teaches the model to generate not only solution rollouts but also the scaffolds that drive those rollouts - jointly optimizing scaffold-rollout co-optimization so the model discovers better search trajectories and produces higher-quality solutions. The architecture is a Qwen3.5-style sparse Mixture-of-Experts with 35B total and 3B active parameters: 40 layers mixing 30 linear attention layers (Gated DeltaNet-style: conv kernel 4, 16 key heads / 32 value heads of dim 128) with 10 full-attention layers (interval 4; 16 Q / 2 KV heads, head dim 256, partial RoPE 0.25, theta 10M, interleaved MRoPE [11,11,10]), 256 routed experts with 8 active per token plus 1 shared expert (expert FFN 512), one Multi-Token Prediction (MTP) layer, 248K vocabulary, a 262,144-token context, and a Qwen3.5 ViT (depth 27, patch 16, spatial merge 2) for image input.
Developed on top of Qwen3.5 and Gemma 4 with additional continued pretraining, mid-training and post-training. Released July 1, 2026 under the MIT License.
Training Data Developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training. Uses scaffold and rollout optimization.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
63.72%
|
19.08.2026 |
|
Terminal-Bench 2.1 (Claude Code)
coding_agent
|
88.51%
|
19.08.2026 |
|
SWE-bench Verified
coding_agent
|
86.24%
|
19.08.2026 |
|
SWE-bench Pro
coding_agent
|
63.00%
|
19.08.2026 |
|
SWE-bench Multilingual
coding_agent
|
76.09%
|
19.08.2026 |
|
DeepSWE
coding_agent
|
0.00
|
19.08.2026 |
|
Frontier-Bench v0.1
coding_agent
|
1.40
|
19.08.2026 |
|
NL2Repo
coding_agent
|
37.38%
|
19.08.2026 |
|
SWE Atlas - QnA
coding_agent
|
50.91%
|
19.08.2026 |
|
Humanity's Last Exam
stem_reasoning
|
35.42%
|
19.08.2026 |
|
HLE (with tools)
stem_reasoning
|
27.12%
|
19.08.2026 |
|
GPQA Diamond
stem_reasoning
|
84.69%
|
19.08.2026 |
|
MCP-Atlas
general_agent
|
70.86%
|
19.08.2026 |
|
Toolathlon Verified
general_agent
|
29.14%
|
19.08.2026 |
|
WideSearch
general_agent
|
55.19%
|
19.08.2026 |
|
BrowseComp
general_agent
|
68.56%
|
19.08.2026 |
|
WildClawBench
coding_agent
|
96.03%
|
19.08.2026 |
Architecture Overview
Architecture Overview
Ornith-1.0-35B is a Mixture-of-Experts (MoE) model built on the qwen3_5_moe architecture.
Key Specifications
| Property | Value |
|---|---|
| Architecture type | Mixture-of-Experts (MoE) |
| Base architecture | qwen3_5_moe |
| Total parameters | 35B (35,000,000,000) |
| Active parameters (MoE) | 3B (3,000,000,000) |
| Tensor type | BF16 |
| Model family | Ornith-1.0 |
Base Models
Post-trained on top of Gemma 4 and Qwen 3.5 with additional continued pretraining, mid-training, and post-training.
Modalities
- Input: text, image
- Output: text
- The model is multimodal (image-text-to-text).
Languages
- English (en)
- Multi-language support
Model Variants in Family
| Variant | Type | Size |
|---|---|---|
| Ornith-1.0-9B | Dense | 9B |
| Ornith-1.0-31B | Dense | 31B |
| Ornith-1.0-35B | MoE | 35B (3B active) |
| Ornith-1.0-397B | MoE | 397B |
Context Length Configuration
Context Length
- Native context length: 262,144 tokens (256K)
- Extended context: Not specified (None in DB)
Serving Configuration
- vLLM:
--max-model-len 262144 - SGLang:
--context-length 262144 - llama.cpp:
-c 262144 - Unsloth:
max_seq_length=262144
Evaluation Context Usage
| Benchmark | Context Used |
|---|---|
| Terminal-Bench 2.1 (Terminus-2) | 128K |
| SWE-Bench Verified/Pro/Multilingual | 256K |
| SWE Atlas | 128K |
| NL2Repo | 400K (extended) |
| ClawEval | 256K |
Self-Improving Training Framework
Self-Improving Training Framework
Ornith-1.0 employs a novel self-improving training framework using reinforcement learning.
Key Concepts
- Joint Optimization of Scaffold and Rollout: Ornith-1.0 uses RL to learn to generate not only solution rollouts but also the scaffold that drives those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.
- Base Models: Post-trained on top of Gemma 4 and Qwen 3.5 with additional continued pretraining, mid-training, and post-training.
- Model Family: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE variants.
Training Stages
- Continued Pretraining (CPT): Additional pretraining on top of base models.
- Mid-training: Intermediate training stage between pretraining and post-training.
- Post-training: Including RL for scaffold and rollout optimization.
- Scaffold and Rollout Optimization: The model learns to generate both the scaffold (search strategy) and the solution (rollout), jointly optimized via RL.
MIT License
License
Ornith-1.0-35B is released under the MIT License.
- MIT licensed, globally accessible, and free from regional limitations.
- Commercial use is allowed.
- The model is fully open-source with no regional restrictions.
Reasoning Model Behavior
Reasoning Model Behavior
Ornith-1.0-35B is a reasoning model: by default the assistant turn opens with a … block before the final answer.
Reasoning Content Separation
When using vLLM or SGLang with a reasoning parser (--reasoning-parser qwen3), the chain-of-thought is returned in a separate reasoning_content field:
reasoning_content: Holds the trace (chain-of-thought)content: Holds the final answer
Tool Call Parsing
The serving recipes enable a tool-call parser so the model's <tool_call> blocks are surfaced as OpenAI-style tool_calls:
- vLLM:
--tool-call-parser qwen3_xml - SGLang:
--tool-call-parser qwen3_coder
Manual Parsing
When using Transformers directly, split on the `` marker:
text = tokenizer.decode(output_ids, skip_special_tokens=True)
if "" in text:
reasoning, answer = text.split("", 1)
reasoning = reasoning.replace("", "").strip()
answer = answer.strip()
else:
reasoning, answer = "", text.strip()
Recommended Sampling Parameters
Recommended Sampling Parameters
Ornith-1.0-35B is a reasoning model. Recommended sampling parameters for optimal performance:
| Parameter | Value | Context |
|---|---|---|
| temperature | 0.6 | General inference, tool calling, agent frameworks |
| top_p | 0.95 | General inference, SWE-Bench evaluations |
| top_k | 20 | Transformers local inference |
| max_new_tokens | 131072 | Terminal-Bench (Claude Code) |
| context_length | 262144 | Native context window (vLLM/SGLang) |
| context_length | 400K | NL2Repo evaluations |
Evaluation-Specific Settings
- SWE-Bench (OpenHands): temp=1.0, top_p=0.95, 256K context
- Terminal-Bench (Terminus-2): temp=1.0, top_p=1.0, 128K context
- Terminal-Bench (Claude Code): temp=1.0, top_p=1.0, max_new_tokens=131072
- SWE Atlas: temp=1.0, top_p=0.95, 128K context
- NL2Repo: temp=1.0, top_p=1.0, 400K context, 48K output
- ClawEval: temp=0.6, 256K context
BibTeX Citation
Citation
If you find our work helpful, feel free to give us a cite.
@misc{ornith-35b,
title = {{Ornith-1.0-35B}: Agentic Coding, Open to All},
url = {https://deep-reinforce.com/ornith_1_0.html},
author = {{DeepReinforce Team}},
year = {2026}
}
Source: Ornith Blog
Agent Framework Integrations
Agent Framework Integrations
Ornith-1.0-35B exposes an OpenAI-compatible endpoint with tool calling, working out of the box with standard agent frameworks.
Supported Frameworks
- Hermes Agent: Point
OPENAI_BASE_URLat your Ornith server, setMODEL=deepreinforce-ai/Ornith-1.0-35B. - Atomic.chat / Ollama / llama.cpp: Load a GGUF build of Ornith.
llama-server -hf deepreinforce-ai/Ornith-1.0-35B-GGUF --port 8000 -c 262144orollama run hf.co/deepreinforce-ai/Ornith-1.0-35B-GGUF. - OpenClaw: Set
OPENAI_MODEL=deepreinforce-ai/Ornith-1.0-35Bwith your Ornith server endpoint. - Unsloth Studio:
pip install unsloth, load withFastLanguageModel.from_pretrained("deepreinforce-ai/Ornith-1.0-35B", max_seq_length=262144, load_in_4bit=True). - OpenHands:
pip install openhands-ai, setLLM_MODEL=openai/deepreinforce-ai/Ornith-1.0-35B, route through LiteLLM. - OpenCode: Register Ornith as a provider in
~/.config/opencode/opencode.jsonwith@ai-sdk/openai-compatible.
MCP Server Example
import os
from openai import OpenAI
client = OpenAI(
base_url=os.getenv("OPENAI_BASE_URL", "http://localhost:8000/v1"),
api_key=os.getenv("OPENAI_API_KEY", "EMPTY"),
)
tools = [{
"type": "function",
"function": {
"name": "run_shell",
"description": "Run a shell command and return its output.",
"parameters": {"type": "object", "properties": {"command": {"type": "string"}}, "required": ["command"]},
},
}]
response = client.chat.completions.create(
model="deepreinforce-ai/Ornith-1.0-35B",
messages=[{"role": "user", "content": "List the Python files in the current directory."}],
tools=tools, temperature=0.6, top_p=0.95,
)
Model Tree
Model Tree for ornith-ai/Ornith-1.0-35B
Ornith-1.0-35B serves as a base model for a wide community ecosystem:
| Type | Count |
|---|---|
| Adapters | 1 model |
| Finetunes | 17 models |
| Merges | 12 models |
| Quantizations | 184 models |
Quantizations are available for multiple runtimes: llama.cpp, LM Studio, Jan, and Ollama.
The model is part of the Ornith-1.0 Collection — a family of open-source LLMs specialized for agentic coding (8 items total).
Local Inference with Transformers
Local Inference with Hugging Face Transformers
For a quick local test or offline generation, load Ornith-1.0-35B directly with Transformers (requires transformers >= 5.8.1).
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "deepreinforce-ai/Ornith-1.0-35B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
generated = model.generate(
**inputs, max_new_tokens=512, do_sample=True,
temperature=0.6, top_p=0.95, top_k=20,
)
output_ids = generated[0][inputs.input_ids.shape[1]:]
content = tokenizer.decode(output_ids, skip_special_tokens=True)
print(content)
Splitting Reasoning from Answer
The reply contains a … reasoning block followed by the final answer. Parse on the `` marker:
text = tokenizer.decode(output_ids, skip_special_tokens=True)
if "" in text:
reasoning, answer = text.split("", 1)
reasoning = reasoning.replace("", "").strip()
answer = answer.strip()
else:
reasoning, answer = "", text.strip()
Chat Completions API Usage
Using Ornith-1.0-35B via the Chat Completions API
Once a vLLM or SGLang server is running, talk to it with any OpenAI-compatible client.
Basic Usage
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Ornith-1.0-35B",
messages=[{"role": "user", "content": "Write a one-line Python lambda that squares a number."}],
temperature=0.6, top_p=0.95, max_tokens=1024,
)
message = response.choices[0].message
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)
Tool Calling
Ornith-1.0-35B emits well-formed function calls that the server parses into the standard tool_calls field.
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
response = client.chat.completions.create(
model="Ornith-1.0-35B",
messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
tools=tools, tool_choice="auto",
temperature=0.6, max_tokens=2048,
)
tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
## -> get_weather {"city": "Paris"}
Any OpenAI-compatible SDK (Python, Node.js) or curl can target the same /v1/chat/completions endpoint.
Serving Recipes
Serving Ornith-1.0-35B
Ornith-1.0-35B requires recent runtimes:
- Transformers ≥ 5.8.1
- vLLM ≥ 0.19.1
- SGLang ≥ 0.5.9
The recipes below stand up an OpenAI-compatible server on a single 8×80GB GPU node (tensor-parallel 8). Adjust --tensor-parallel-size / --tp to the number of GPUs available.
vLLM
vllm serve deepreinforce-ai/Ornith-1.0-35B \
--served-model-name Ornith-1.0-35B \
--tensor-parallel-size 8 \
--host 0.0.0.0 --port 8000 \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--trust-remote-code
SGLang
python -m sglang.launch_server \
--model-path deepreinforce-ai/Ornith-1.0-35B \
--served-model-name Ornith-1.0-35B \
--tp 8 \
--host 0.0.0.0 --port 8000 \
--context-length 262144 \
--mem-fraction-static 0.85 \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3
Hugging Face Transformers
Requires transformers >= 5.8.1. See the Transformers installation guide.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "deepreinforce-ai/Ornith-1.0-35B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto", device_map="auto")
Agentic Coding Benchmark Results
Agentic Coding Benchmark Results
| Benchmark | Ornith-1.0-35B | Qwen3.5-35B | Qwen3.6-35B | Gemma4-31B | Qwen3.5-397B |
|---|---|---|---|---|---|
| Agentic Coding | |||||
| Terminal-Bench 2.1 (Terminus-2) | 64.2 | 41.4 | 52.5 | 42.1 | 53.5 |
| Terminal-Bench 2.1 (Claude Code) | 62.8 | 38.9 | 49.2 | - | 48.6 |
| SWE-bench Verified | 75.6 | 70 | 73.4 | 52 | 76.4 |
| SWE-bench Pro | 50.4 | 44.6 | 49.5 | 35.7 | 51.6 |
| SWE-bench Multilingual | 69.3 | 60.3 | 67.2 | 51.7 | 69.3 |
| NL2Repo | 34.6 | 20.5 | 29.4 | 15.5 | 36.8 |
| Claw-eval Avg | 69.8 | 65.4 | 68.7 | 48.5 | 70.7 |
| SWE Atlas - QnA | 37.1 | 13.2 | 15.5 | - | 20.4 |
| SWE Atlas - RF | 29.7 | 10.2 | 11.4 | - | 18.4 |
| SWE Atlas - TW | 27.8 | 9.8 | 13.3 | - | 18.5 |
Evaluation Methodology
- Terminal-Bench 2.1 (Terminus-2): Harbor/Terminus-2 framework, parser=json, temperature=1.0, top_p=1.0, 128K context, 4h timeout, 32 CPU cores, 48GB RAM, avg of 5 runs.
- Terminal-Bench 2.1 (Claude Code): Claude Code 2.1.126, parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072, avg of 5 runs.
- SWE-Bench Verified/Pro/Multilingual: OpenHands harness, temp=1.0, top_p=0.95, 256k context window.
- SWE Atlas QnA/RF/TW: mini SWE agent harness, temp=1.0, top_p=0.95, 128K context, avg of 5 runs.
- NL2Repo: temperature=1.0, top_p=1.0, 400K context, 48K output, anti-hacking filters.
- ClawEval: Agentic code benchmark over real-user task distributions, temp=0.6, 256K context.
Key Highlights
Ornith-1.0 Key Highlights
Ornith-1.0 is a self-improving family of open-source models for agentic coding.
- State-of-the-Art Coding Agents: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
- Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scaffold that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.
- Licence: MIT licensed, globally accessible, and free from regional limitations.
The 35B-MoE variant is the lightweight member of the Ornith family, designed for efficient single-GPU deployment.
Architecture
- Attention
- Hybrid Attention (16:2)
- MoE
- 256 experts · top-8 per token
- Layers
- 40
- Hidden size
- 2048
- Context
- 262K tokens
- RoPE θ
- 10M
- Parameters
- 35000M
- Active params
- 3000M
Source: extracted model record · model repo
Training Pipeline
-
1
cpt
Continued pretraining on top of Qwen3.5 and Gemma 4
Additional continued pretraining, mid-training and post-training on top of Gemma 4 and Qwen 3.5 bases (DB training_data_info).
-
2
other
Mid-training
Mid-training stage between CPT and post-training (DB training_data_info).
-
3
rl
Self-improving scaffold-rollout RL (post-training)
RL jointly optimizing scaffolds and solution rollouts; discovers better search trajectories, higher-quality solutions (card Self-Improving Training Framework).
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| Agentic coding RL rollouts (solution + scaffold) | rl | — | — |
Linked Resources
deepreinforce-ai (GitHub)
https://github.com/deepreinforce-ai
GGUF builds (deepreinforce-ai/Ornith-1.0-35B, llama.cpp/Ollama serve)
https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B
Ornith (HuggingFace collection)
https://huggingface.co/ornith-ai
Hermes agent runtime (OpenAI-compatible endpoint)
https://huggingface.co/ornith-ai/Ornith-1.0-35B
OpenClaw / OpenHands / opencode agent frameworks usage
https://huggingface.co/ornith-ai/Ornith-1.0-35B#agentic-usage
Citation (Ornith-1.0)
https://huggingface.co/ornith-ai/Ornith-1.0-35B#citation
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 282,300 pulls | daily | 01.09.2026 |
| ollama | downloads | 262,300 pulls | daily | 31.08.2026 |
| ollama | downloads | 198,400 pulls | daily | 30.08.2026 |
| ollama | downloads | 142,200 pulls | daily | 29.08.2026 |
| ollama | downloads | 117,900 pulls | daily | 28.08.2026 |
| ollama | downloads | 65,100 pulls | daily | 27.08.2026 |
| ollama | downloads | 44,800 pulls | daily | 26.08.2026 |
| ollama | downloads | 36,900 pulls | daily | 25.08.2026 |
| ollama | downloads | 31,000 pulls | daily | 24.08.2026 |