Parameters
9.0B
Architecture
Dense decoder-only (post-trained on Qwen 3.5 9B), reasoning model with tool-calling
Released
21.06.2026
License
MIT License
Input Modalities
Output Modalities
Context (native)
262,144 tokens
Context (extended)
262,144 tokens
About
Ornith-1.0-9B (ornith-ai/Ornith-1.0-9B) is the most lightweight member of the Ornith-1.0 family - a self-improving, dense 9B-parameter (post-trained on top of the Qwen 3.5 9B base) open-source model for agentic coding, designed for efficient single-GPU deployment. It is MIT licensed, globally accessible, and free from regional limitations.
Ornith-1.0's scaffold-rollout co-optimization training framework uses RL to learn to generate not only solution rollouts but also the scaffolds that drive those rollouts: by jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions. The 9B model achieves state-of-the-art results among open-source models of comparable size on agentic coding benchmarks (Terminal-Bench 2.1, SWE-bench Verified/Pro/Multilingual, NL2Repo, Claw-eval, SWE Atlas), outperforming its Qwen3.5-9B base (e.g. 69.4 vs 53.2 on SWE-bench Verified). It natively emits imdi ... dimdi reasoning blocks and supports tool calling; GGUF builds are published at deepreinforce-ai for llama.cpp, Ollama, OpenClaw, OpenHands, Hermes and opencode integration. Context length 262,144 tokens (native; no extension documented).
Training Data Post-trained (RL, self-scaffolding) on top of Qwen 3.5 9B base. RL jointly optimizes scaffold and solution rollouts for agentic coding.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
32.60%
|
25.06.2026 |
|
Terminal-Bench 2.1 (Claude Code)
coding_agent
|
43.75%
|
25.06.2026 |
|
SWE-bench Verified
coding_agent
|
79.13%
|
25.06.2026 |
|
SWE-bench Pro
coding_agent
|
53.62%
|
25.06.2026 |
|
SWE-bench Multilingual
coding_agent
|
55.62%
|
25.06.2026 |
|
NL2Repo
coding_agent
|
26.00%
|
25.06.2026 |
|
Claw-Eval Avg
coding_agent
|
75.78%
|
25.06.2026 |
|
SWE Atlas - QnA
coding_agent
|
15.88%
|
25.06.2026 |
|
SWE Atlas - RF
coding_agent
|
22.08%
|
25.06.2026 |
|
SWE Atlas - TW
coding_agent
|
16.95%
|
25.06.2026 |
Parsing Reasoning from Answer
Parsing Reasoning from Answer
When using HuggingFace Transformers directly (without a serving runtime), split the reasoning trace from the final answer by parsing on the </think> marker:
text = tokenizer.decode(output_ids, skip_special_tokens=True)
if "</think>" in text:
reasoning, answer = text.split("</think>", 1)
reasoning = reasoning.replace("<think>", "").strip()
answer = answer.strip()
else:
reasoning, answer = "", text.strip()
This separates the model's internal reasoning (the <think> block) from the final user-facing answer.
Reasoning Mode and Thinking Blocks
Reasoning Mode and Thinking Blocks
Ornith-1.0-9B is a reasoning model: by default the assistant turn opens with a <think> … </think> block before the final answer.
Serving with Reasoning Parser
When served via vLLM or SGLang with --reasoning-parser qwen3, the chain-of-thought is returned in a separate reasoning_content field:
message = response.choices[0].message
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)
Tool Call Parser
Ornith-1.0-9B also supports tool calling via the qwen3_xml (vLLM) or qwen3_coder (SGLang) tool-call parser, which surfaces <tool_call> blocks as OpenAI-style tool_calls.
License Information
License Information
License: MIT
Ornith-1.0-9B is MIT licensed, globally accessible, and free from regional limitations.
- Repository and weights: MIT licensed
- Commercial use: Allowed
- Open source: Yes
- Regional restrictions: None
This makes Ornith-1.0-9B suitable for both research and commercial applications without licensing constraints.
Model Tree and Community
Model Tree and Community
Model Tree for ornith-ai/Ornith-1.0-9B
| Type | Count |
|---|---|
| Adapters | 2 models |
| Finetunes | 35 models |
| Quantizations | 105 models |
Quantizations are available for: llama.cpp, LM Studio, Jan, and Ollama.
Collection
Ornith-1.0 — A family of open-source LLMs specialized for agentic coding.
- 8 items in the collection
- Updated 5 days ago
- 388 views
Community
- Spaces using this model: 10
- Followers: 2,850
- Downloads (last month): 2,206,037
- Likes: 538
- Community discussions: 15
Base Model
This model is listed as a base model for finetunes, quantizations, and adapters. The Ornith-1.0 family includes:
- 9B-Dense (this model)
- 31B-Dense
- 35B-MoE
- 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5)
BibTeX Citation
Citation
If you find this work helpful, please cite:
@misc{ornith_9b,
title = {{Ornith-1.0-9B}: Agentic Coding, Open to All},
url = {https://deep-reinforce.com/ornith_1_0.html},
author = {{DeepReinforce Team}},
year = {2026}
}
- Blog post: https://deep-reinforce.com/ornith.html
- Ornith 1.0 blog: https://deep-reinforce.com/ornith_1_0.html
Recommended Sampling Parameters
Recommended Sampling Parameters
For optimal performance with Ornith-1.0-9B:
| Parameter | Value | Notes |
|---|---|---|
| temperature | 0.6 | Default for general use |
| top_p | 0.95 | Nucleus sampling |
| top_k | 20 | Top-k sampling |
Benchmark Reproduction
To reproduce the reported benchmark scores, use:
| Parameter | Value |
|---|---|
| temperature | 1.0 |
| top_p | 1.0 |
⚠️ Note: Benchmark evaluation uses temperature=1.0 and top_p=1.0, while general usage recommends temperature=0.6 for more focused outputs.
Development Tool and Framework Integrations
Integration with Development Tools and Frameworks
Ornith-1.0-9B is optimized for terminal-based coding agents and works with any OpenAI-compatible endpoint.
Hermes Agent
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export MODEL="deepreinforce-ai/Ornith-1.0-9B"
Atomic.chat / Ollama / llama.cpp
## llama.cpp — serve an OpenAI-compatible API on port 8000
llama-server -hf deepreinforce-ai/Ornith-1.0-9B-GGUF --port 8000 -c 262144
## Ollama — pull and chat with the same GGUF straight from Hugging Face
ollama run hf.co/deepreinforce-ai/Ornith-1.0-9B-GGUF
OpenClaw
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export OPENAI_MODEL="deepreinforce-ai/Ornith-1.0-9B"
Unsloth Studio
pip install unsloth
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
"deepreinforce-ai/Ornith-1.0-9B",
max_seq_length=262144,
load_in_4bit=True,
)
OpenHands
pip install openhands-ai
export LLM_MODEL="openai/deepreinforce-ai/Ornith-1.0-9B"
export LLM_BASE_URL="http://localhost:8000/v1"
export LLM_API_KEY="EMPTY"
openhands
OpenCode
// ~/.config/opencode/opencode.json
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"ornith": {
"npm": "@ai-sdk/openai-compatible",
"name": "Ornith (local)",
"options": { "baseURL": "http://localhost:8000/v1", "apiKey": "EMPTY" },
"models": { "deepreinforce-ai/Ornith-1.0-9B": { "name": "Ornith-1.0-9B" } }
}
}
}
Agent Framework Integration
Agent Framework Integration via MCP
Ornith-1.0-9B exposes an OpenAI-compatible endpoint with tool calling, so it works out of the box with standard agent frameworks. Below is a minimal example connecting Ornith to tools through an MCP server:
import os
from openai import OpenAI
client = OpenAI(
base_url=os.getenv("OPENAI_BASE_URL", "http://localhost:8000/v1"),
api_key=os.getenv("OPENAI_API_KEY", "EMPTY"),
)
tools = [
{
"type": "function",
"function": {
"name": "run_shell",
"description": "Run a shell command and return its output.",
"parameters": {
"type": "object",
"properties": {
"command": {"type": "string", "description": "The command to run"}
},
"required": ["command"],
},
},
}
]
messages = [{"role": "user", "content": "List the Python files in the current directory."}]
response = client.chat.completions.create(
model="deepreinforce-ai/Ornith-1.0-9B",
messages=messages,
tools=tools,
temperature=0.6,
top_p=0.95,
)
print(response.choices[0].message)
Tool Calling via Chat Completions API
Tool Calling via Chat Completions API
Ornith-1.0-9B emits well-formed function calls that the serving runtime parses into the standard tool_calls field:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY",
)
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
]
response = client.chat.completions.create(
model="Ornith-1.0-9B",
messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
tools=tools,
tool_choice="auto",
temperature=0.6,
max_tokens=2048,
)
tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
## -> get_weather {"city": "Paris"}
Any OpenAI-compatible SDK (Python, Node.js, etc.) or curl can target the same endpoint.
Chat Completions API – Basic Usage
Basic Chat Completions API Usage
Once a vLLM or SGLang server is running, talk to it with any OpenAI-compatible client:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY", # any non-empty string works for a local server
)
response = client.chat.completions.create(
model="Ornith-1.0-9B",
messages=[
{"role": "user", "content": "Write a one-line Python lambda that squares a number."}
],
temperature=0.6,
top_p=0.95,
max_tokens=1024,
)
message = response.choices[0].message
## reasoning_content holds the <think> trace; content holds the final answer.
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)
You can also stream tokens, or use curl with the same /v1/chat/completions endpoint.
HuggingFace Transformers Inference
Serving with Hugging Face Transformers
For quick local testing or offline generation, load the model directly with Transformers (≥ 5.8.1):
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "deepreinforce-ai/Ornith-1.0-9B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
dtype="auto",
device_map="auto",
)
messages = [
{"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
generated = model.generate(
**inputs,
max_new_tokens=512,
do_sample=True,
temperature=0.6,
top_p=0.95,
top_k=20,
)
output_ids = generated[0][inputs.input_ids.shape[1]:]
## The reply contains a <think> ... </think> reasoning block followed by the answer.
content = tokenizer.decode(output_ids, skip_special_tokens=True)
print(content)
See the Transformers installation guide for setup instructions.
SGLang Serving
Serving with SGLang
Ornith-1.0-9B can be served via SGLang with an OpenAI-compatible API:
python -m sglang.launch_server \
--model-path deepreinforce-ai/Ornith-1.0-9B \
--served-model-name Ornith-1.0-9B \
--host 0.0.0.0 --port 8000 \
--context-length 262144 \
--mem-fraction-static 0.85 \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3
Key flags:
--reasoning-parser qwen3— separates thinking blocks intoreasoning_contentfield--tool-call-parser qwen3_coder— parses tool calls into standard format--context-length 262144— full 256K context window--mem-fraction-static 0.85— static memory fraction for KV cache
Minimum version: SGLang ≥ 0.5.9
vLLM Serving
Serving with vLLM
Ornith-1.0-9B serves comfortably on a single 80GB GPU via vLLM with an OpenAI-compatible server:
vllm serve deepreinforce-ai/Ornith-1.0-9B \
--served-model-name Ornith-1.0-9B \
--host 0.0.0.0 --port 8000 \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--trust-remote-code
Key flags:
--reasoning-parser qwen3— separates thinking blocks intoreasoning_contentfield--tool-call-parser qwen3_xml— parses tool calls into OpenAI-styletool_calls--enable-prefix-caching— optimizes repeated prompt prefixes--max-model-len 262144— full 256K context window
Add --tensor-parallel-size N to shard across multiple GPUs.
Runtime Requirements
Runtime Requirements
Serving Ornith-1.0-9B requires recent runtime versions:
| Runtime | Minimum Version |
|---|---|
| Transformers | ≥ 5.8.1 |
| vLLM | ≥ 0.19.1 |
| SGLang | ≥ 0.5.9 |
Hardware Requirements
- Model size in memory: ~19 GB in bf16
- Recommended GPU: Single 80GB GPU (serves comfortably)
- Optional:
--tensor-parallel-size/--tpfor multi-GPU sharding
Model Tags
qwen3_5— based on Qwen 3.5 architectureimage-text-to-text— multimodal input supportconversational— chat interface supportsafetensors— distributed in Safetensors format
Agentic Coding Benchmarks
Agentic Coding Benchmark Results
| Benchmark | Ornith-1.0-9B | Qwen3.5-9B | Qwen3.5-35B | Gemma4-12B | Gemma4-31B |
|---|---|---|---|---|---|
| Agentic Coding | |||||
| Terminal-Bench 2.1 (Terminus-2) | 43.1 | 21.3 | 41.4 | 21 | 42.1 |
| Terminal-Bench 2.1 (Claude Code) | 40.6 | 18.9 | 38.9 | – | – |
| SWE-bench Verified | 69.4 | 53.2 | 70 | 44.2 | 52 |
| SWE-bench Pro | 42.9 | 31.3 | 44.6 | 27.6 | 35.7 |
| SWE-bench Multilingual | 52.0 | 39.7 | 60.3 | 32.5 | 51.7 |
| NL2Repo | 27.2 | 16.2 | 20.5 | 10.3 | 15.5 |
| Claw-eval Avg | 63.1 | 53.2 | 65.4 | 32.5 | 48.5 |
| SWE Atlas - QnA | 17.9 | 9.2 | 13.2 | – | – |
| SWE Atlas - RF | 16.6 | 4.3 | 10.2 | – | – |
| SWE Atlas - TW | 15.3 | 4.4 | 9.8 | – | – |
Evaluation Setup
- Terminal-Bench 2.1 (Terminus-2): Harbor/Terminus-2 framework, parser=json, temperature=1.0, top_p=1.0, 128K context, 4h timeout, 32 CPU cores, 48GB RAM, avg of 5 runs. Qwen chat template adjusted for consistency.
- Terminal-Bench 2.1 (Claude Code): Claude Code 2.1.126, parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072, avg of 5 runs.
- SWE-bench Verified/Pro/Multilingual: OpenHands harness, temp=1.0, top_p=0.95, 256K context window.
- SWE Atlas QnA/RF/TW: mini SWE agent harness, temp=1.0, top_p=0.95, 128K context, avg of 5 runs.
- NL2Repo: temp=1.0, top_p=1.0, 400K context, 48K output, anti-hacking filters.
- ClawEval: Agentic code benchmark over real-user task distributions, temp=0.6, 256K context.
Source: HuggingFace Model Card
Model Architecture
Model Architecture
Architecture Type: Dense decoder-only (post-trained on Qwen 3.5 9B)
| Property | Value |
|---|---|
| Base Model | Qwen 3.5 9B |
| Parameters | ~9B (900M active in safetensors) |
| Architecture Tag | qwen3_5 |
| Pipeline Tag | text-generation, image-text-to-text |
| Libraries | Transformers, Safetensors |
| Tensor Type | BF16 |
| Model Size in Memory | ~19 GB in bf16 |
Key Capabilities:
- Reasoning model (generates thinking blocks before final answer)
- Tool-calling support (emits well-formed function calls)
- Multimodal: accepts text and image inputs, outputs text
- Conversational interface support
Post-Training: RL-based post-training on Qwen 3.5 9B base, jointly optimizing scaffold and solution rollouts for agentic coding.
Key Highlights
Key Highlights
Ornith-1.0-9B is part of the Ornith-1.0 family — a self-improving family of open-source models for agentic coding.
State-of-the-Art Coding Agents
Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
Self-Improving Training Framework
Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scaffold that drives those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.
MIT License
MIT licensed, globally accessible, and free from regional limitations.
Lightweight Design
Ornith-1.0-9B is the most lightweight member of the Ornith family, designed for efficient single-GPU deployment.
Architecture
- Attention
- Hybrid Attention (16:4)
- Layers
- 32
- Hidden size
- 4096
- Context
- 262K tokens
- Parameters
- 9000M
Source: Hugging Face config.json · Qwen3_5ForConditionalGeneration · exact layer pattern · model repo
rl_scaffold
Yes
Yes
Training Pipeline
-
1
rl
Self-Scaffolding RL for Agentic Coding
Ornith-1.0 employs RL to jointly optimize the scaffold and the resulting solution rollouts, enabling the model to discover better search trajectories and generate higher-quality solutions. Post-trained on top of Qwen 3.5 9B.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| Scaffold + solution RL rollouts (agentic coding) | rl | — | — |
Linked Resources
Ornith-1 GitHub Repository
https://github.com/deepreinforce-ai/Ornith-1
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding
https://deep-reinforce.com/ornith_1_0.html
Ornith Blog
https://deep-reinforce.com/ornith.html
Ornith-1.0 Collection
https://huggingface.co/collections/deepreinforce-ai/ornith-10
BibTeX citation for Ornith-1.0-9B
https://deep-reinforce.com/ornith_1_0.html
Trend Analysis
24h Change
+0.9%
7d Change
+9.3%
Current
3,185
downloads
-2.3%
downloads_all_time
+1.0%
downloads
+0.8%
likes
+0.0%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 443,200 pulls | daily | 01.09.2026 |
| huggingface | downloads_all_time | 4,164,857 | daily | 01.09.2026 |
| huggingface | followers | 3,185 | daily | 01.09.2026 |
| huggingface | likes | 540 | daily | 01.09.2026 |
| huggingface | downloads | 1,764,324 | daily | 01.09.2026 |
| ollama | downloads | 439,500 pulls | daily | 31.08.2026 |
| huggingface | downloads_all_time | 4,125,350 | daily | 31.08.2026 |
| huggingface | followers | 3,158 | daily | 31.08.2026 |
| huggingface | likes | 540 | daily | 31.08.2026 |
| huggingface | downloads | 1,805,743 | daily | 31.08.2026 |
| ollama | downloads | 435,700 pulls | daily | 30.08.2026 |
| huggingface | downloads_all_time | 4,097,118 | daily | 30.08.2026 |
| huggingface | followers | 3,129 | daily | 30.08.2026 |
| huggingface | likes | 540 | daily | 30.08.2026 |
| huggingface | downloads | 1,838,530 | daily | 30.08.2026 |
| ollama | downloads | 432,100 pulls | daily | 29.08.2026 |
| huggingface | followers | 3,096 | daily | 29.08.2026 |
| huggingface | likes | 539 | daily | 29.08.2026 |
| huggingface | downloads | 1,855,431 | daily | 29.08.2026 |
| ollama | downloads | 428,200 pulls | daily | 28.08.2026 |
| huggingface | followers | 3,061 | daily | 28.08.2026 |
| huggingface | likes | 538 | daily | 28.08.2026 |
| huggingface | downloads | 1,893,740 | daily | 28.08.2026 |
| ollama | downloads | 424,600 pulls | daily | 27.08.2026 |
| huggingface | followers | 3,020 | daily | 27.08.2026 |
| huggingface | likes | 538 | daily | 27.08.2026 |
| huggingface | downloads | 2,015,459 | daily | 27.08.2026 |
| ollama | downloads | 421,400 pulls | daily | 26.08.2026 |
| huggingface | followers | 2,976 | daily | 26.08.2026 |
| huggingface | likes | 539 | daily | 26.08.2026 |
| huggingface | downloads | 2,124,015 | daily | 26.08.2026 |
| ollama | downloads | 418,400 pulls | daily | 25.08.2026 |
| huggingface | followers | 2,913 | daily | 25.08.2026 |
| huggingface | likes | 539 | daily | 25.08.2026 |
| huggingface | downloads | 2,170,130 | daily | 25.08.2026 |
| ollama | downloads | 415,200 pulls | daily | 24.08.2026 |
| huggingface | followers | 2,860 | daily | 24.08.2026 |
| huggingface | likes | 538 | daily | 24.08.2026 |
| huggingface | downloads | 2,206,037 | daily | 24.08.2026 |
| huggingface | followers | 2,803 | daily | 23.08.2026 |
| huggingface | likes | 538 | daily | 23.08.2026 |
| huggingface | downloads | 2,150,443 | daily | 23.08.2026 |
| huggingface | followers | 2,741 | daily | 22.08.2026 |
| huggingface | likes | 538 | daily | 22.08.2026 |
| huggingface | downloads | 2,184,736 | daily | 22.08.2026 |
| huggingface | followers | 2,673 | daily | 21.08.2026 |
| huggingface | likes | 535 | daily | 21.08.2026 |
| huggingface | downloads | 2,258,799 | daily | 21.08.2026 |
| huggingface | followers | 2,560 | daily | 20.08.2026 |
| huggingface | likes | 535 | daily | 20.08.2026 |