Ornith-1.0-9B

DeepReinforce AI

Parameters

9.0B

Architecture

Dense decoder-only (post-trained on Qwen 3.5 9B), reasoning model with tool-calling

Released

21.06.2026

License

MIT License

Open Weights Commercial Use Multimodal BF16 Ornith en

Input Modalities

text image

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

262,144 tokens

Openness Index Score 100.0/100

About

Ornith-1.0-9B (ornith-ai/Ornith-1.0-9B) is the most lightweight member of the Ornith-1.0 family - a self-improving, dense 9B-parameter (post-trained on top of the Qwen 3.5 9B base) open-source model for agentic coding, designed for efficient single-GPU deployment. It is MIT licensed, globally accessible, and free from regional limitations.

Ornith-1.0's scaffold-rollout co-optimization training framework uses RL to learn to generate not only solution rollouts but also the scaffolds that drive those rollouts: by jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions. The 9B model achieves state-of-the-art results among open-source models of comparable size on agentic coding benchmarks (Terminal-Bench 2.1, SWE-bench Verified/Pro/Multilingual, NL2Repo, Claw-eval, SWE Atlas), outperforming its Qwen3.5-9B base (e.g. 69.4 vs 53.2 on SWE-bench Verified). It natively emits imdi ... dimdi reasoning blocks and supports tool calling; GGUF builds are published at deepreinforce-ai for llama.cpp, Ollama, OpenClaw, OpenHands, Hermes and opencode integration. Context length 262,144 tokens (native; no extension documented).

Training Data Post-trained (RL, self-scaffolding) on top of Qwen 3.5 9B base. RL jointly optimizes scaffold and solution rollouts for agentic coding.

Benchmark Scores

Benchmark Score Date
Terminal-Bench 2.1 (Terminus-2)
coding_agent
32.60%
25.06.2026
Terminal-Bench 2.1 (Claude Code)
coding_agent
43.75%
25.06.2026
SWE-bench Verified
coding_agent
79.13%
25.06.2026
SWE-bench Pro
coding_agent
53.62%
25.06.2026
SWE-bench Multilingual
coding_agent
55.62%
25.06.2026
NL2Repo
coding_agent
26.00%
25.06.2026
Claw-Eval Avg
coding_agent
75.78%
25.06.2026
SWE Atlas - QnA
coding_agent
15.88%
25.06.2026
SWE Atlas - RF
coding_agent
22.08%
25.06.2026
SWE Atlas - TW
coding_agent
16.95%
25.06.2026

Parsing Reasoning from Answer

Parsing Reasoning from Answer

When using HuggingFace Transformers directly (without a serving runtime), split the reasoning trace from the final answer by parsing on the </think> marker:

text = tokenizer.decode(output_ids, skip_special_tokens=True)
if "</think>" in text:
    reasoning, answer = text.split("</think>", 1)
    reasoning = reasoning.replace("<think>", "").strip()
    answer = answer.strip()
else:
    reasoning, answer = "", text.strip()

This separates the model's internal reasoning (the <think> block) from the final user-facing answer.

Reasoning Mode and Thinking Blocks

Reasoning Mode and Thinking Blocks

Ornith-1.0-9B is a reasoning model: by default the assistant turn opens with a <think> … </think> block before the final answer.

Serving with Reasoning Parser

When served via vLLM or SGLang with --reasoning-parser qwen3, the chain-of-thought is returned in a separate reasoning_content field:

message = response.choices[0].message
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)

Tool Call Parser

Ornith-1.0-9B also supports tool calling via the qwen3_xml (vLLM) or qwen3_coder (SGLang) tool-call parser, which surfaces <tool_call> blocks as OpenAI-style tool_calls.

License Information

License Information

License: MIT

Ornith-1.0-9B is MIT licensed, globally accessible, and free from regional limitations.

  • Repository and weights: MIT licensed
  • Commercial use: Allowed
  • Open source: Yes
  • Regional restrictions: None

This makes Ornith-1.0-9B suitable for both research and commercial applications without licensing constraints.

Model Tree and Community

Model Tree and Community

Model Tree for ornith-ai/Ornith-1.0-9B

Type Count
Adapters 2 models
Finetunes 35 models
Quantizations 105 models

Quantizations are available for: llama.cpp, LM Studio, Jan, and Ollama.

Collection

Ornith-1.0 — A family of open-source LLMs specialized for agentic coding.

  • 8 items in the collection
  • Updated 5 days ago
  • 388 views

Community

  • Spaces using this model: 10
  • Followers: 2,850
  • Downloads (last month): 2,206,037
  • Likes: 538
  • Community discussions: 15

Base Model

This model is listed as a base model for finetunes, quantizations, and adapters. The Ornith-1.0 family includes:

  • 9B-Dense (this model)
  • 31B-Dense
  • 35B-MoE
  • 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5)

BibTeX Citation

Citation

If you find this work helpful, please cite:

@misc{ornith_9b,
    title = {{Ornith-1.0-9B}: Agentic Coding, Open to All},
    url = {https://deep-reinforce.com/ornith_1_0.html},
    author = {{DeepReinforce Team}},
    year = {2026}
}
  • Blog post: https://deep-reinforce.com/ornith.html
  • Ornith 1.0 blog: https://deep-reinforce.com/ornith_1_0.html

Recommended Sampling Parameters

Recommended Sampling Parameters

For optimal performance with Ornith-1.0-9B:

Parameter Value Notes
temperature 0.6 Default for general use
top_p 0.95 Nucleus sampling
top_k 20 Top-k sampling

Benchmark Reproduction

To reproduce the reported benchmark scores, use:

Parameter Value
temperature 1.0
top_p 1.0

⚠️ Note: Benchmark evaluation uses temperature=1.0 and top_p=1.0, while general usage recommends temperature=0.6 for more focused outputs.

Development Tool and Framework Integrations

Integration with Development Tools and Frameworks

Ornith-1.0-9B is optimized for terminal-based coding agents and works with any OpenAI-compatible endpoint.

Hermes Agent

export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export MODEL="deepreinforce-ai/Ornith-1.0-9B"

Atomic.chat / Ollama / llama.cpp

## llama.cpp — serve an OpenAI-compatible API on port 8000
llama-server -hf deepreinforce-ai/Ornith-1.0-9B-GGUF --port 8000 -c 262144

## Ollama — pull and chat with the same GGUF straight from Hugging Face
ollama run hf.co/deepreinforce-ai/Ornith-1.0-9B-GGUF

OpenClaw

export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export OPENAI_MODEL="deepreinforce-ai/Ornith-1.0-9B"

Unsloth Studio

pip install unsloth

from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
    "deepreinforce-ai/Ornith-1.0-9B",
    max_seq_length=262144,
    load_in_4bit=True,
)

OpenHands

pip install openhands-ai
export LLM_MODEL="openai/deepreinforce-ai/Ornith-1.0-9B"
export LLM_BASE_URL="http://localhost:8000/v1"
export LLM_API_KEY="EMPTY"
openhands

OpenCode

// ~/.config/opencode/opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "ornith": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Ornith (local)",
      "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "EMPTY" },
      "models": { "deepreinforce-ai/Ornith-1.0-9B": { "name": "Ornith-1.0-9B" } }
    }
  }
}

Agent Framework Integration

Agent Framework Integration via MCP

Ornith-1.0-9B exposes an OpenAI-compatible endpoint with tool calling, so it works out of the box with standard agent frameworks. Below is a minimal example connecting Ornith to tools through an MCP server:

import os
from openai import OpenAI

client = OpenAI(
    base_url=os.getenv("OPENAI_BASE_URL", "http://localhost:8000/v1"),
    api_key=os.getenv("OPENAI_API_KEY", "EMPTY"),
)

tools = [
    {
        "type": "function",
        "function": {
            "name": "run_shell",
            "description": "Run a shell command and return its output.",
            "parameters": {
                "type": "object",
                "properties": {
                    "command": {"type": "string", "description": "The command to run"}
                },
                "required": ["command"],
            },
        },
    }
]

messages = [{"role": "user", "content": "List the Python files in the current directory."}]

response = client.chat.completions.create(
    model="deepreinforce-ai/Ornith-1.0-9B",
    messages=messages,
    tools=tools,
    temperature=0.6,
    top_p=0.95,
)
print(response.choices[0].message)

Tool Calling via Chat Completions API

Tool Calling via Chat Completions API

Ornith-1.0-9B emits well-formed function calls that the serving runtime parses into the standard tool_calls field:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="EMPTY",
)

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a city",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
                "required": ["city"],
            },
        },
    }
]

response = client.chat.completions.create(
    model="Ornith-1.0-9B",
    messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
    tools=tools,
    tool_choice="auto",
    temperature=0.6,
    max_tokens=2048,
)

tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
## -> get_weather {"city": "Paris"}

Any OpenAI-compatible SDK (Python, Node.js, etc.) or curl can target the same endpoint.

Chat Completions API – Basic Usage

Basic Chat Completions API Usage

Once a vLLM or SGLang server is running, talk to it with any OpenAI-compatible client:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="EMPTY",  # any non-empty string works for a local server
)

response = client.chat.completions.create(
    model="Ornith-1.0-9B",
    messages=[
        {"role": "user", "content": "Write a one-line Python lambda that squares a number."}
    ],
    temperature=0.6,
    top_p=0.95,
    max_tokens=1024,
)

message = response.choices[0].message
## reasoning_content holds the <think> trace; content holds the final answer.
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)

You can also stream tokens, or use curl with the same /v1/chat/completions endpoint.

HuggingFace Transformers Inference

Serving with Hugging Face Transformers

For quick local testing or offline generation, load the model directly with Transformers (≥ 5.8.1):

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "deepreinforce-ai/Ornith-1.0-9B"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    dtype="auto",
    device_map="auto",
)

messages = [
    {"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

inputs = tokenizer(text, return_tensors="pt").to(model.device)
generated = model.generate(
    **inputs,
    max_new_tokens=512,
    do_sample=True,
    temperature=0.6,
    top_p=0.95,
    top_k=20,
)
output_ids = generated[0][inputs.input_ids.shape[1]:]

## The reply contains a <think> ... </think> reasoning block followed by the answer.
content = tokenizer.decode(output_ids, skip_special_tokens=True)
print(content)

See the Transformers installation guide for setup instructions.

SGLang Serving

Serving with SGLang

Ornith-1.0-9B can be served via SGLang with an OpenAI-compatible API:

python -m sglang.launch_server \
    --model-path deepreinforce-ai/Ornith-1.0-9B \
    --served-model-name Ornith-1.0-9B \
    --host 0.0.0.0 --port 8000 \
    --context-length 262144 \
    --mem-fraction-static 0.85 \
    --tool-call-parser qwen3_coder \
    --reasoning-parser qwen3

Key flags:

  • --reasoning-parser qwen3 — separates thinking blocks into reasoning_content field
  • --tool-call-parser qwen3_coder — parses tool calls into standard format
  • --context-length 262144 — full 256K context window
  • --mem-fraction-static 0.85 — static memory fraction for KV cache

Minimum version: SGLang ≥ 0.5.9

vLLM Serving

Serving with vLLM

Ornith-1.0-9B serves comfortably on a single 80GB GPU via vLLM with an OpenAI-compatible server:

vllm serve deepreinforce-ai/Ornith-1.0-9B \
    --served-model-name Ornith-1.0-9B \
    --host 0.0.0.0 --port 8000 \
    --max-model-len 262144 \
    --gpu-memory-utilization 0.90 \
    --enable-prefix-caching \
    --enable-auto-tool-choice --tool-call-parser qwen3_xml \
    --reasoning-parser qwen3 \
    --trust-remote-code

Key flags:

  • --reasoning-parser qwen3 — separates thinking blocks into reasoning_content field
  • --tool-call-parser qwen3_xml — parses tool calls into OpenAI-style tool_calls
  • --enable-prefix-caching — optimizes repeated prompt prefixes
  • --max-model-len 262144 — full 256K context window

Add --tensor-parallel-size N to shard across multiple GPUs.

Runtime Requirements

Runtime Requirements

Serving Ornith-1.0-9B requires recent runtime versions:

Runtime Minimum Version
Transformers ≥ 5.8.1
vLLM ≥ 0.19.1
SGLang ≥ 0.5.9

Hardware Requirements

  • Model size in memory: ~19 GB in bf16
  • Recommended GPU: Single 80GB GPU (serves comfortably)
  • Optional: --tensor-parallel-size / --tp for multi-GPU sharding

Model Tags

  • qwen3_5 — based on Qwen 3.5 architecture
  • image-text-to-text — multimodal input support
  • conversational — chat interface support
  • safetensors — distributed in Safetensors format

Agentic Coding Benchmarks

Agentic Coding Benchmark Results

Benchmark Ornith-1.0-9B Qwen3.5-9B Qwen3.5-35B Gemma4-12B Gemma4-31B
Agentic Coding
Terminal-Bench 2.1 (Terminus-2) 43.1 21.3 41.4 21 42.1
Terminal-Bench 2.1 (Claude Code) 40.6 18.9 38.9 – –
SWE-bench Verified 69.4 53.2 70 44.2 52
SWE-bench Pro 42.9 31.3 44.6 27.6 35.7
SWE-bench Multilingual 52.0 39.7 60.3 32.5 51.7
NL2Repo 27.2 16.2 20.5 10.3 15.5
Claw-eval Avg 63.1 53.2 65.4 32.5 48.5
SWE Atlas - QnA 17.9 9.2 13.2 – –
SWE Atlas - RF 16.6 4.3 10.2 – –
SWE Atlas - TW 15.3 4.4 9.8 – –

Evaluation Setup

  • Terminal-Bench 2.1 (Terminus-2): Harbor/Terminus-2 framework, parser=json, temperature=1.0, top_p=1.0, 128K context, 4h timeout, 32 CPU cores, 48GB RAM, avg of 5 runs. Qwen chat template adjusted for consistency.
  • Terminal-Bench 2.1 (Claude Code): Claude Code 2.1.126, parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072, avg of 5 runs.
  • SWE-bench Verified/Pro/Multilingual: OpenHands harness, temp=1.0, top_p=0.95, 256K context window.
  • SWE Atlas QnA/RF/TW: mini SWE agent harness, temp=1.0, top_p=0.95, 128K context, avg of 5 runs.
  • NL2Repo: temp=1.0, top_p=1.0, 400K context, 48K output, anti-hacking filters.
  • ClawEval: Agentic code benchmark over real-user task distributions, temp=0.6, 256K context.

Source: HuggingFace Model Card

Model Architecture

Model Architecture

Architecture Type: Dense decoder-only (post-trained on Qwen 3.5 9B)

Property Value
Base Model Qwen 3.5 9B
Parameters ~9B (900M active in safetensors)
Architecture Tag qwen3_5
Pipeline Tag text-generation, image-text-to-text
Libraries Transformers, Safetensors
Tensor Type BF16
Model Size in Memory ~19 GB in bf16

Key Capabilities:

  • Reasoning model (generates thinking blocks before final answer)
  • Tool-calling support (emits well-formed function calls)
  • Multimodal: accepts text and image inputs, outputs text
  • Conversational interface support

Post-Training: RL-based post-training on Qwen 3.5 9B base, jointly optimizing scaffold and solution rollouts for agentic coding.

Key Highlights

Key Highlights

Ornith-1.0-9B is part of the Ornith-1.0 family — a self-improving family of open-source models for agentic coding.

State-of-the-Art Coding Agents

Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.

Self-Improving Training Framework

Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scaffold that drives those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.

MIT License

MIT licensed, globally accessible, and free from regional limitations.

Lightweight Design

Ornith-1.0-9B is the most lightweight member of the Ornith family, designed for efficient single-GPU deployment.

Architecture

Decoder Block ×32 input Embedding vocab 248K · d 4096 Linear / Recurrent Hybrid 16:4 · dₕ 256 ×24 Full Attention Hybrid 16:4 · dₕ 256 ×8 Dense FFN silu · d 12K Final Norm LM Head vocab 248K output
Attention
Hybrid Attention (16:4)
Layers
32
Hidden size
4096
Context
262K tokens
Parameters
9000M

Source: Hugging Face config.json · Qwen3_5ForConditionalGeneration · exact layer pattern · model repo

Type: Dense decoder-only Transformer (Qwen3.5-based)
Attention: Grouped-query attention (Qwen3.5)
Decoder: autoregressive
Total parameters 9
Context length 262K
Vision Yes
Post Training

rl_scaffold

Reasoning

Yes

Tool Calling

Yes

Training Pipeline

  1. 1
    rl

    Self-Scaffolding RL for Agentic Coding

    Ornith-1.0 employs RL to jointly optimize the scaffold and the resulting solution rollouts, enabling the model to discover better search trajectories and generate higher-quality solutions. Post-trained on top of Qwen 3.5 9B.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Scaffold + solution RL rollouts (agentic coding) rl — —

Trend Analysis

24h Change

+0.9%

7d Change

+9.3%

Current

3,185

huggingface

downloads

-2.3%

huggingface

downloads_all_time

+1.0%

ollama

downloads

+0.8%

huggingface

likes

+0.0%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 443,200 pulls daily 01.09.2026
huggingface downloads_all_time 4,164,857 daily 01.09.2026
huggingface followers 3,185 daily 01.09.2026
huggingface likes 540 daily 01.09.2026
huggingface downloads 1,764,324 daily 01.09.2026
ollama downloads 439,500 pulls daily 31.08.2026
huggingface downloads_all_time 4,125,350 daily 31.08.2026
huggingface followers 3,158 daily 31.08.2026
huggingface likes 540 daily 31.08.2026
huggingface downloads 1,805,743 daily 31.08.2026
ollama downloads 435,700 pulls daily 30.08.2026
huggingface downloads_all_time 4,097,118 daily 30.08.2026
huggingface followers 3,129 daily 30.08.2026
huggingface likes 540 daily 30.08.2026
huggingface downloads 1,838,530 daily 30.08.2026
ollama downloads 432,100 pulls daily 29.08.2026
huggingface followers 3,096 daily 29.08.2026
huggingface likes 539 daily 29.08.2026
huggingface downloads 1,855,431 daily 29.08.2026
ollama downloads 428,200 pulls daily 28.08.2026
huggingface followers 3,061 daily 28.08.2026
huggingface likes 538 daily 28.08.2026
huggingface downloads 1,893,740 daily 28.08.2026
ollama downloads 424,600 pulls daily 27.08.2026
huggingface followers 3,020 daily 27.08.2026
huggingface likes 538 daily 27.08.2026
huggingface downloads 2,015,459 daily 27.08.2026
ollama downloads 421,400 pulls daily 26.08.2026
huggingface followers 2,976 daily 26.08.2026
huggingface likes 539 daily 26.08.2026
huggingface downloads 2,124,015 daily 26.08.2026
ollama downloads 418,400 pulls daily 25.08.2026
huggingface followers 2,913 daily 25.08.2026
huggingface likes 539 daily 25.08.2026
huggingface downloads 2,170,130 daily 25.08.2026
ollama downloads 415,200 pulls daily 24.08.2026
huggingface followers 2,860 daily 24.08.2026
huggingface likes 538 daily 24.08.2026
huggingface downloads 2,206,037 daily 24.08.2026
huggingface followers 2,803 daily 23.08.2026
huggingface likes 538 daily 23.08.2026
huggingface downloads 2,150,443 daily 23.08.2026
huggingface followers 2,741 daily 22.08.2026
huggingface likes 538 daily 22.08.2026
huggingface downloads 2,184,736 daily 22.08.2026
huggingface followers 2,673 daily 21.08.2026
huggingface likes 535 daily 21.08.2026
huggingface downloads 2,258,799 daily 21.08.2026
huggingface followers 2,560 daily 20.08.2026
huggingface likes 535 daily 20.08.2026

View full metric history →

Related Models