Parameters
35.0B total / 3.0B active
MoE: total / active
Architecture
MoE decoder-only (qwen35moe), post-trained on Qwen 3.5; reasoning model with tool-calling. 35B total / ~3B active (256 experts, 8 routed + 1 shared).
Released
21.06.2026
License
MIT License
Input Modalities
Output Modalities
Context (native)
262,144 tokens
Context (extended)
262,144 tokens
About
Ornith-1.0-35B (GGUF builds at deepreinforce-ai/Ornith-1.0-35B; card at ornith-ai/Ornith-1.0-35B) is the lightweight 35B-MoE member of Ornith 1.0, the self-improving family of open-source models for agentic coding from deepreinforce-ai, designed for efficient single-GPU deployment. Ornith 1.0 ships in 9B-Dense, 31B-Dense, 35B-MoE and 397B-MoE variants post-trained on top of Gemma 4 and Qwen 3.5, achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
Its signature is the self-improving training framework: RL teaches the model to generate not only solution rollouts but also the scaffolds that drive them - jointly optimizing scaffold-rollout co-optimization so the model discovers better search trajectories and higher-quality solutions. This record's lineage: post-trained with RL (self-scaffolding) on top of the Qwen3.5 35B base. The underlying model is a Qwen3.5-style sparse Mixture-of-Experts with 35B total and 3B active parameters: 40 layers mixing 30 linear-attention layers (Gated DeltaNet-style) with 10 full-attention layers (interval 4), 256 routed experts with 8 active per token plus 1 shared expert, one Multi-Token Prediction layer, 248K vocabulary and a 262,144-token context, with an optional Qwen3.5 ViT for image input on the multimodal builds.
Serving is optimized for GGUF runtimes: llama.cpp, Ollama, LM Studio and vLLM/SGLang OpenAI-compatible endpoints, with agent frameworks (Hermes, OpenClaw, OpenHands, opencode) and Unsloth for local fine-tuning. Released June 21, 2026 under the MIT License.
Training Data Post-trained (RL with self-scaffolding) on top of Qwen 3.5 35B base. RL jointly optimizes scaffold and solution rollouts for agentic coding.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
63.72%
|
25.06.2026 |
|
Terminal-Bench 2.1 (Claude Code)
coding_agent
|
88.51%
|
25.06.2026 |
|
SWE-bench Verified
coding_agent
|
86.24%
|
25.06.2026 |
|
SWE-bench Pro
coding_agent
|
63.00%
|
25.06.2026 |
|
SWE-bench Multilingual
coding_agent
|
76.09%
|
25.06.2026 |
|
NL2Repo
coding_agent
|
37.38%
|
25.06.2026 |
|
Claw-Eval Avg
coding_agent
|
84.19%
|
25.06.2026 |
|
SWE Atlas - QnA
coding_agent
|
50.91%
|
25.06.2026 |
|
SWE Atlas - RF
coding_agent
|
45.60%
|
25.06.2026 |
|
SWE Atlas - TW
coding_agent
|
36.39%
|
25.06.2026 |
Model Tree and Spaces
Model tree for ornith-ai/Ornith-1.0-35B
Adapters
Finetunes
Merges
Quantizations
Spaces using ornith-ai/Ornith-1.0-35B 3
Collection including ornith-ai/Ornith-1.0-35B
[
Ornith-1.0
Collection
Ornith-1.0 is a family of open-source LLMs specialized for agentic coding. • 8 items • Updated 23 days ago • 394
](https://huggingface.co/collections/ornith-ai/ornith-10)
Citation
Agentic Usage: agent frameworks (Hermes, OpenClaw, OpenHands, opencode, Unsloth)
Agentic Usage
Ornith-1.0-35B excels in tool-calling and agentic coding capabilities.
Agent Frameworks
Because Ornith-1.0-35B exposes an OpenAI-compatible endpoint with tool calling, it works out of the box with standard agent frameworks. Below is a minimal example that connects Ornith-1.0-35B to tools through an MCP server.
import os
from openai import OpenAI
client = OpenAI(
base_url=os.getenv("OPENAI_BASE_URL", "http://localhost:8000/v1"),
api_key=os.getenv("OPENAI_API_KEY", "EMPTY"),
)
tools = [
{
"type": "function",
"function": {
"name": "run_shell",
"description": "Run a shell command and return its output.",
"parameters": {
"type": "object",
"properties": {
"command": {"type": "string", "description": "The command to run"}
},
"required": ["command"],
},
},
}
]
messages = [{"role": "user", "content": "List the Python files in the current directory."}]
response = client.chat.completions.create(
model="deepreinforce-ai/Ornith-1.0-35B",
messages=messages,
tools=tools,
temperature=0.6,
top_p=0.95,
)
print(response.choices[0].message)
Examples of using Ornith with agent harness:
Hermes Agent
## Hermes talks to any OpenAI-compatible endpoint — point it at your Ornith server.
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export MODEL="deepreinforce-ai/Ornith-1.0-35B"
Atomic.chat/ Ollama / llama.cpp
## Both runtimes load a GGUF build of Ornith (publish one at deepreinforce-ai/Ornith-1.0-35B-GGUF).
## llama.cpp — serve an OpenAI-compatible API on port 8000.
llama-server -hf deepreinforce-ai/Ornith-1.0-35B-GGUF --port 8000 -c 262144
## Ollama — pull and chat with the same GGUF straight from Hugging Face.
ollama run hf.co/deepreinforce-ai/Ornith-1.0-35B-GGUF
OpenClaw
## OpenClaw talks to any OpenAI-compatible endpoint — point it at your Ornith server.
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export OPENAI_MODEL="deepreinforce-ai/Ornith-1.0-35B"
Unsloth Studio
pip install unsloth
## Load Ornith for fast local inference or fine-tuning (Python):
## from unsloth import FastLanguageModel
## model, tokenizer = FastLanguageModel.from_pretrained(
## "deepreinforce-ai/Ornith-1.0-35B",
## max_seq_length=262144,
## load_in_4bit=True,
## )
OpenHands
pip install openhands-ai
## OpenHands routes through LiteLLM; the "openai/" prefix selects the OpenAI-compatible path.
export LLM_MODEL="openai/deepreinforce-ai/Ornith-1.0-35B"
export LLM_BASE_URL="http://localhost:8000/v1"
export LLM_API_KEY="EMPTY"
## Launch the CLI (or run the official OpenHands Docker image with the same env vars).
openhands
Coding CLIs
Ornith-1.0-35B is optimized for terminal-based coding agents. Point any OpenAI-compatible coding CLI at your Ornith-1.0-35B endpoint (set OPENAI_BASE_URL and OPENAI_API_KEY) to understand large codebases, automate tedious work, and ship faster.
OpenCode
## Register your local Ornith endpoint as a provider in ~/.config/opencode/opencode.json:
##
## {
## "$schema": "https://opencode.ai/config.json",
## "provider": {
## "ornith": {
## "npm": "@ai-sdk/openai-compatible",
## "name": "Ornith (local)",
## "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "EMPTY" },
## "models": { "deepreinforce-ai/Ornith-1.0-35B": { "name": "Ornith-1.0-35B" } }
## }
## }
## }
opencode
Citation
If you find our work helpful, feel free to give us a cite.
@misc{ornith-35b,
title = {{Ornith-1.0-35B}: Agentic Coding, Open to All},
url = {https://deep-reinforce.com/ornith_1_0.html},
author = {{DeepReinforce Team}},
year = {2026}
}
Model size
665k params
Tensor type
BF16
·
Serving Ornith-1.0-35B (reasoning content, tool calls)
Serving Ornith-1.0-35B
The two recipes below stand up an OpenAI-compatible server on a single 8×80GB GPU node (tensor-parallel 8). Adjust --tensor-parallel-size / --tp to the number of GPUs you have.
vLLM
vllm serve deepreinforce-ai/Ornith-1.0-35B \
--served-model-name Ornith-1.0-35B \
--tensor-parallel-size 8 \
--host 0.0.0.0 --port 8000 \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--trust-remote-code
SGLang
python -m sglang.launch_server \
--model-path deepreinforce-ai/Ornith-1.0-35B \
--served-model-name Ornith-1.0-35B \
--tp 8 \
--host 0.0.0.0 --port 8000 \
--context-length 262144 \
--mem-fraction-static 0.85 \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3
Hugging Face Transformers
For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-35B requires transformers >= 5.8.1.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "deepreinforce-ai/Ornith-1.0-35B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
dtype="auto",
device_map="auto",
)
messages = [
{"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
generated = model.generate(
**inputs,
max_new_tokens=512,
do_sample=True,
temperature=0.6,
top_p=0.95,
top_k=20,
)
output_ids = generated[0][inputs.input_ids.shape[1]:]
## The reply contains a <think> ... </think> reasoning block followed by the answer.
content = tokenizer.decode(output_ids, skip_special_tokens=True)
print(content)
To split the reasoning trace from the final answer, parse on the </think> marker:
text = tokenizer.decode(output_ids, skip_special_tokens=True)
if "</think>" in text:
reasoning, answer = text.split("</think>", 1)
reasoning = reasoning.replace("<think>", "").strip()
answer = answer.strip()
else:
reasoning, answer = "", text.strip()
Quickstart
Quickstart
📝 NOTE
Ornith-1.0-35B is a reasoning model: by default the assistant turn opens with a <think> … </think> block before the final answer. The serving recipes below enable a reasoning parser so the chain-of-thought is returned in a separate reasoning_content field, and a tool-call parser so the model's <tool_call> blocks are surfaced as OpenAI-style tool_calls.
Serving Ornith-1.0-35B requires recent runtimes:
- Transformers ≥ 5.8.1
- vLLM ≥ 0.19.1
- SGLang ≥ 0.5.9
Benchmarks (vs Qwen3.5-35B, Qwen3.6-35B, Gemma4-31B, Qwen3.5-397B)
Benchmarks
| Ornith-1.0-35B | Qwen3.5-35B | Qwen3.6-35B | Gemma4-31B | Qwen3.5-397B | |
|---|---|---|---|---|---|
| Agentic Coding | |||||
| Terminal-Bench 2.1 (Terminus-2) | 64.2 | 41.4 | 52.5 | 42.1 | 53.5 |
| Terminal-Bench 2.1 (Claude Code) | 62.8 | 38.9 | 49.2 | - | 48.6 |
| SWE-bench Verified | 75.6 | 70 | 73.4 | 52 | 76.4 |
| SWE-bench Pro | 50.4 | 44.6 | 49.5 | 35.7 | 51.6 |
| SWE-bench Multilingual | 69.3 | 60.3 | 67.2 | 51.7 | 69.3 |
| NL2Repo | 34.6 | 20.5 | 29.4 | 15.5 | 36.8 |
| Claw-eval Avg | 69.8 | 65.4 | 68.7 | 48.5 | 70.7 |
| SWE Atlas - QnA | 37.1 | 13.2 | 15.5 | - | 20.4 |
| SWE Atlas - RF | 29.7 | 10.2 | 11.4 | - | 18.4 |
| SWE Atlas - TW | 27.8 | 9.8 | 13.3 | - | 18.5 |
* Terminal-Bench 2.1 (Terminus-2): We evaluate Terminal-Bench 2.1 using the Harbor/Terminus-2 framework with parser=json, temperature=1.0, top_p=1.0, and a 128K context window. Each run uses a 4-hour timeout with 32 CPU cores and 48GB RAM, and results are averaged over 5 runs. We adjust the Qwen chat template to ensure consistency between training and inference (https://huggingface.co/deepreinforce-ai/Ornith-1.0-397B/blob/main/chat_template.jinja), and modify Harbor to align with vLLM's reasoning_content key.
* Terminal-Bench 2.1 (Claude Code): We evaluate Terminal-Bench 2.1 using Claude Code 2.1.126 with parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072. Results are averaged over 5 runs. Again, Qwen chat template needs to be modified.
* SWE-Bench Verified, Pro and Multilingual: using OpenHands harness with temp=1.0, top_p=0.95, 256k context window.
* SWE Atlas QnA, RF, TW: using mini SWE agent harness with temp=1.0, top_p=0.95, 128K context window. Results are averaged over 5 runs.
* NL2Repo: with temperature=1.0, top_p=1.0, 400K context, 48K output and anti-hacking filters.
* ClawEval: An agentic code benchmark over real-user task distributions; temp=0.6 and 256K context.
Ornith-1.0-35B Introduction (self-improving agentic coding)
Ornith-1.0-35B
Aloha! 🌺 Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding.
Highlights:
- State-of-the-Art Coding Agents: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
- Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scallfold that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.
- Licence: MIT licensed, globally accessible, and free from regional limitations.

Ornith 1.0 35B
This model card documents Ornith-1.0-35B, the lightweight member of the Ornith family, designed for efficient single-GPU deployment.
Architecture
- Attention
- Hybrid Attention (16:2)
- MoE
- 256 experts · top-8 per token
- Layers
- 40
- Hidden size
- 2048
- Context
- 262K tokens
- Parameters
- 35000M
- Active params
- 3000M
Source: Hugging Face config.json · Qwen3_5MoeForConditionalGeneration · exact layer pattern · model repo
Training Pipeline
-
1
rl
RL self-scaffolding post-training on Qwen 3.5 35B
Ornith-1.0 employs RL to jointly optimize scaffold and solution rollouts for agentic coding. Post-trained on top of Qwen 3.5 35B base.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| Agentic coding RL rollouts (solution + scaffold) | rl | — | — |
Linked Resources
Ornith Blog (DeepReinforce)
https://deep-reinforce.com/ornith.html
deepreinforce-ai/Ornith-1
https://github.com/deepreinforce-ai/Ornith-1
Ornith-1.0 Collection
https://huggingface.co/collections/deepreinforce-ai/ornith-10
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding
https://deep-reinforce.com/ornith_1_0.html
Ornith-1.0-35B Citation (BibTeX)
https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B-GGUF
Trend Analysis
24h Change
+0.9%
7d Change
+9.3%
Current
3,185
downloads
+0.4%
downloads_all_time
+1.2%
likes
+0.0%
downloads
+0.1%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 3,239 pulls | daily | 01.09.2026 |
| huggingface | downloads_all_time | 6,075,393 | daily | 01.09.2026 |
| huggingface | followers | 3,185 | daily | 01.09.2026 |
| huggingface | likes | 504 | daily | 01.09.2026 |
| huggingface | downloads | 3,056,578 | daily | 01.09.2026 |
| ollama | downloads | 3,236 pulls | daily | 31.08.2026 |
| huggingface | downloads_all_time | 6,003,553 | daily | 31.08.2026 |
| huggingface | followers | 3,158 | daily | 31.08.2026 |
| huggingface | likes | 504 | daily | 31.08.2026 |
| huggingface | downloads | 3,045,441 | daily | 31.08.2026 |
| ollama | downloads | 3,231 pulls | daily | 30.08.2026 |
| huggingface | downloads_all_time | 5,952,173 | daily | 30.08.2026 |
| huggingface | followers | 3,129 | daily | 30.08.2026 |
| huggingface | likes | 504 | daily | 30.08.2026 |
| huggingface | downloads | 3,034,579 | daily | 30.08.2026 |
| ollama | downloads | 3,227 pulls | daily | 29.08.2026 |
| huggingface | followers | 3,096 | daily | 29.08.2026 |
| huggingface | likes | 504 | daily | 29.08.2026 |
| huggingface | downloads | 2,994,126 | daily | 29.08.2026 |
| ollama | downloads | 3,218 pulls | daily | 28.08.2026 |
| huggingface | followers | 3,061 | daily | 28.08.2026 |
| huggingface | likes | 503 | daily | 28.08.2026 |
| huggingface | downloads | 3,005,458 | daily | 28.08.2026 |
| ollama | downloads | 3,207 pulls | daily | 27.08.2026 |
| huggingface | followers | 3,020 | daily | 27.08.2026 |
| huggingface | likes | 503 | daily | 27.08.2026 |
| huggingface | downloads | 3,127,649 | daily | 27.08.2026 |
| ollama | downloads | 3,197 pulls | daily | 26.08.2026 |
| huggingface | followers | 2,976 | daily | 26.08.2026 |
| huggingface | likes | 504 | daily | 26.08.2026 |
| huggingface | downloads | 3,234,380 | daily | 26.08.2026 |
| ollama | downloads | 3,186 pulls | daily | 25.08.2026 |
| huggingface | followers | 2,913 | daily | 25.08.2026 |
| huggingface | likes | 505 | daily | 25.08.2026 |
| huggingface | downloads | 3,286,288 | daily | 25.08.2026 |
| ollama | downloads | 3,181 pulls | daily | 24.08.2026 |
| huggingface | followers | 2,860 | daily | 24.08.2026 |
| huggingface | likes | 505 | daily | 24.08.2026 |
| huggingface | downloads | 3,323,741 | daily | 24.08.2026 |
| huggingface | followers | 2,803 | daily | 23.08.2026 |
| huggingface | likes | 505 | daily | 23.08.2026 |
| huggingface | downloads | 3,188,828 | daily | 23.08.2026 |
| huggingface | followers | 2,741 | daily | 22.08.2026 |
| huggingface | likes | 505 | daily | 22.08.2026 |
| huggingface | downloads | 3,192,730 | daily | 22.08.2026 |
| huggingface | followers | 2,673 | daily | 21.08.2026 |
| huggingface | likes | 505 | daily | 21.08.2026 |
| huggingface | downloads | 3,228,856 | daily | 21.08.2026 |
| huggingface | followers | 2,560 | daily | 20.08.2026 |
| huggingface | likes | 504 | daily | 20.08.2026 |