Ornith-1.0-35B-A3B

Ornith AI

Parameters

35.0B total / 3.0B active

MoE: total / active

Architecture

Mixture-of-Experts (qwen3_5_moe)

Released

01.07.2026

License

MIT License

Open Weights Commercial Use Multimodal BF16 Ornith-1.0 en multi

Input Modalities

text image

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

262,144 tokens

Openness Index Score 100.0/100

About

Ornith-1.0-35B (ornith-ai/Ornith-1.0-35B) is the lightweight 35B MoE member of Ornith 1.0, deepreinforce-ai's self-improving family of open-source models for agentic coding, designed for efficient single-GPU deployment. Ornith 1.0 ships in 9B-Dense, 31B-Dense, 35B-MoE and 397B-MoE variants post-trained on top of Gemma 4 and Qwen 3.5, achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.

Its signature is the self-improving training framework: reinforcement learning teaches the model to generate not only solution rollouts but also the scaffolds that drive those rollouts - jointly optimizing scaffold-rollout co-optimization so the model discovers better search trajectories and produces higher-quality solutions. The architecture is a Qwen3.5-style sparse Mixture-of-Experts with 35B total and 3B active parameters: 40 layers mixing 30 linear attention layers (Gated DeltaNet-style: conv kernel 4, 16 key heads / 32 value heads of dim 128) with 10 full-attention layers (interval 4; 16 Q / 2 KV heads, head dim 256, partial RoPE 0.25, theta 10M, interleaved MRoPE [11,11,10]), 256 routed experts with 8 active per token plus 1 shared expert (expert FFN 512), one Multi-Token Prediction (MTP) layer, 248K vocabulary, a 262,144-token context, and a Qwen3.5 ViT (depth 27, patch 16, spatial merge 2) for image input.

Developed on top of Qwen3.5 and Gemma 4 with additional continued pretraining, mid-training and post-training. Released July 1, 2026 under the MIT License.

Training Data Developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training. Uses scaffold and rollout optimization.

Benchmark Scores

Benchmark Score Date
Terminal-Bench 2.1 (Terminus-2)
coding_agent
63.72%
19.08.2026
Terminal-Bench 2.1 (Claude Code)
coding_agent
88.51%
19.08.2026
SWE-bench Verified
coding_agent
86.24%
19.08.2026
SWE-bench Pro
coding_agent
63.00%
19.08.2026
SWE-bench Multilingual
coding_agent
76.09%
19.08.2026
DeepSWE
coding_agent
0.00
19.08.2026
Frontier-Bench v0.1
coding_agent
1.40
19.08.2026
NL2Repo
coding_agent
37.38%
19.08.2026
SWE Atlas - QnA
coding_agent
50.91%
19.08.2026
Humanity's Last Exam
stem_reasoning
35.42%
19.08.2026
HLE (with tools)
stem_reasoning
27.12%
19.08.2026
GPQA Diamond
stem_reasoning
84.69%
19.08.2026
MCP-Atlas
general_agent
70.86%
19.08.2026
Toolathlon Verified
general_agent
29.14%
19.08.2026
WideSearch
general_agent
55.19%
19.08.2026
BrowseComp
general_agent
68.56%
19.08.2026
WildClawBench
coding_agent
96.03%
19.08.2026

Architecture Overview

Architecture Overview

Ornith-1.0-35B is a Mixture-of-Experts (MoE) model built on the qwen3_5_moe architecture.

Key Specifications

Property Value
Architecture type Mixture-of-Experts (MoE)
Base architecture qwen3_5_moe
Total parameters 35B (35,000,000,000)
Active parameters (MoE) 3B (3,000,000,000)
Tensor type BF16
Model family Ornith-1.0

Base Models

Post-trained on top of Gemma 4 and Qwen 3.5 with additional continued pretraining, mid-training, and post-training.

Modalities

  • Input: text, image
  • Output: text
  • The model is multimodal (image-text-to-text).

Languages

  • English (en)
  • Multi-language support

Model Variants in Family

Variant Type Size
Ornith-1.0-9B Dense 9B
Ornith-1.0-31B Dense 31B
Ornith-1.0-35B MoE 35B (3B active)
Ornith-1.0-397B MoE 397B

Context Length Configuration

Context Length

  • Native context length: 262,144 tokens (256K)
  • Extended context: Not specified (None in DB)

Serving Configuration

  • vLLM: --max-model-len 262144
  • SGLang: --context-length 262144
  • llama.cpp: -c 262144
  • Unsloth: max_seq_length=262144

Evaluation Context Usage

Benchmark Context Used
Terminal-Bench 2.1 (Terminus-2) 128K
SWE-Bench Verified/Pro/Multilingual 256K
SWE Atlas 128K
NL2Repo 400K (extended)
ClawEval 256K

Self-Improving Training Framework

Self-Improving Training Framework

Ornith-1.0 employs a novel self-improving training framework using reinforcement learning.

Key Concepts

  • Joint Optimization of Scaffold and Rollout: Ornith-1.0 uses RL to learn to generate not only solution rollouts but also the scaffold that drives those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.
  • Base Models: Post-trained on top of Gemma 4 and Qwen 3.5 with additional continued pretraining, mid-training, and post-training.
  • Model Family: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE variants.

Training Stages

  1. Continued Pretraining (CPT): Additional pretraining on top of base models.
  2. Mid-training: Intermediate training stage between pretraining and post-training.
  3. Post-training: Including RL for scaffold and rollout optimization.
  4. Scaffold and Rollout Optimization: The model learns to generate both the scaffold (search strategy) and the solution (rollout), jointly optimized via RL.

MIT License

License

Ornith-1.0-35B is released under the MIT License.

  • MIT licensed, globally accessible, and free from regional limitations.
  • Commercial use is allowed.
  • The model is fully open-source with no regional restrictions.

Reasoning Model Behavior

Reasoning Model Behavior

Ornith-1.0-35B is a reasoning model: by default the assistant turn opens with a … block before the final answer.

Reasoning Content Separation

When using vLLM or SGLang with a reasoning parser (--reasoning-parser qwen3), the chain-of-thought is returned in a separate reasoning_content field:

  • reasoning_content: Holds the trace (chain-of-thought)
  • content: Holds the final answer

Tool Call Parsing

The serving recipes enable a tool-call parser so the model's <tool_call> blocks are surfaced as OpenAI-style tool_calls:

  • vLLM: --tool-call-parser qwen3_xml
  • SGLang: --tool-call-parser qwen3_coder

Manual Parsing

When using Transformers directly, split on the `` marker:

text = tokenizer.decode(output_ids, skip_special_tokens=True)
if "" in text:
    reasoning, answer = text.split("", 1)
    reasoning = reasoning.replace("", "").strip()
    answer = answer.strip()
else:
    reasoning, answer = "", text.strip()

Recommended Sampling Parameters

Recommended Sampling Parameters

Ornith-1.0-35B is a reasoning model. Recommended sampling parameters for optimal performance:

Parameter Value Context
temperature 0.6 General inference, tool calling, agent frameworks
top_p 0.95 General inference, SWE-Bench evaluations
top_k 20 Transformers local inference
max_new_tokens 131072 Terminal-Bench (Claude Code)
context_length 262144 Native context window (vLLM/SGLang)
context_length 400K NL2Repo evaluations

Evaluation-Specific Settings

  • SWE-Bench (OpenHands): temp=1.0, top_p=0.95, 256K context
  • Terminal-Bench (Terminus-2): temp=1.0, top_p=1.0, 128K context
  • Terminal-Bench (Claude Code): temp=1.0, top_p=1.0, max_new_tokens=131072
  • SWE Atlas: temp=1.0, top_p=0.95, 128K context
  • NL2Repo: temp=1.0, top_p=1.0, 400K context, 48K output
  • ClawEval: temp=0.6, 256K context

BibTeX Citation

Citation

If you find our work helpful, feel free to give us a cite.

@misc{ornith-35b,
    title = {{Ornith-1.0-35B}: Agentic Coding, Open to All},
    url = {https://deep-reinforce.com/ornith_1_0.html},
    author = {{DeepReinforce Team}},
    year = {2026}
}

Source: Ornith Blog

Agent Framework Integrations

Agent Framework Integrations

Ornith-1.0-35B exposes an OpenAI-compatible endpoint with tool calling, working out of the box with standard agent frameworks.

Supported Frameworks

  • Hermes Agent: Point OPENAI_BASE_URL at your Ornith server, set MODEL=deepreinforce-ai/Ornith-1.0-35B.
  • Atomic.chat / Ollama / llama.cpp: Load a GGUF build of Ornith. llama-server -hf deepreinforce-ai/Ornith-1.0-35B-GGUF --port 8000 -c 262144 or ollama run hf.co/deepreinforce-ai/Ornith-1.0-35B-GGUF.
  • OpenClaw: Set OPENAI_MODEL=deepreinforce-ai/Ornith-1.0-35B with your Ornith server endpoint.
  • Unsloth Studio: pip install unsloth, load with FastLanguageModel.from_pretrained("deepreinforce-ai/Ornith-1.0-35B", max_seq_length=262144, load_in_4bit=True).
  • OpenHands: pip install openhands-ai, set LLM_MODEL=openai/deepreinforce-ai/Ornith-1.0-35B, route through LiteLLM.
  • OpenCode: Register Ornith as a provider in ~/.config/opencode/opencode.json with @ai-sdk/openai-compatible.

MCP Server Example

import os
from openai import OpenAI

client = OpenAI(
    base_url=os.getenv("OPENAI_BASE_URL", "http://localhost:8000/v1"),
    api_key=os.getenv("OPENAI_API_KEY", "EMPTY"),
)
tools = [{
    "type": "function",
    "function": {
        "name": "run_shell",
        "description": "Run a shell command and return its output.",
        "parameters": {"type": "object", "properties": {"command": {"type": "string"}}, "required": ["command"]},
    },
}]
response = client.chat.completions.create(
    model="deepreinforce-ai/Ornith-1.0-35B",
    messages=[{"role": "user", "content": "List the Python files in the current directory."}],
    tools=tools, temperature=0.6, top_p=0.95,
)

Model Tree

Model Tree for ornith-ai/Ornith-1.0-35B

Ornith-1.0-35B serves as a base model for a wide community ecosystem:

Type Count
Adapters 1 model
Finetunes 17 models
Merges 12 models
Quantizations 184 models

Quantizations are available for multiple runtimes: llama.cpp, LM Studio, Jan, and Ollama.

The model is part of the Ornith-1.0 Collection — a family of open-source LLMs specialized for agentic coding (8 items total).

Local Inference with Transformers

Local Inference with Hugging Face Transformers

For a quick local test or offline generation, load Ornith-1.0-35B directly with Transformers (requires transformers >= 5.8.1).

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "deepreinforce-ai/Ornith-1.0-35B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto", device_map="auto")

messages = [{"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
generated = model.generate(
    **inputs, max_new_tokens=512, do_sample=True,
    temperature=0.6, top_p=0.95, top_k=20,
)
output_ids = generated[0][inputs.input_ids.shape[1]:]
content = tokenizer.decode(output_ids, skip_special_tokens=True)
print(content)

Splitting Reasoning from Answer

The reply contains a … reasoning block followed by the final answer. Parse on the `` marker:

text = tokenizer.decode(output_ids, skip_special_tokens=True)
if "" in text:
    reasoning, answer = text.split("", 1)
    reasoning = reasoning.replace("", "").strip()
    answer = answer.strip()
else:
    reasoning, answer = "", text.strip()

Chat Completions API Usage

Using Ornith-1.0-35B via the Chat Completions API

Once a vLLM or SGLang server is running, talk to it with any OpenAI-compatible client.

Basic Usage

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="Ornith-1.0-35B",
    messages=[{"role": "user", "content": "Write a one-line Python lambda that squares a number."}],
    temperature=0.6, top_p=0.95, max_tokens=1024,
)

message = response.choices[0].message
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)

Tool Calling

Ornith-1.0-35B emits well-formed function calls that the server parses into the standard tool_calls field.

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

response = client.chat.completions.create(
    model="Ornith-1.0-35B",
    messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
    tools=tools, tool_choice="auto",
    temperature=0.6, max_tokens=2048,
)

tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
## -> get_weather {"city": "Paris"}

Any OpenAI-compatible SDK (Python, Node.js) or curl can target the same /v1/chat/completions endpoint.

Serving Recipes

Serving Ornith-1.0-35B

Ornith-1.0-35B requires recent runtimes:

  • Transformers ≥ 5.8.1
  • vLLM ≥ 0.19.1
  • SGLang ≥ 0.5.9

The recipes below stand up an OpenAI-compatible server on a single 8×80GB GPU node (tensor-parallel 8). Adjust --tensor-parallel-size / --tp to the number of GPUs available.

vLLM

vllm serve deepreinforce-ai/Ornith-1.0-35B \
    --served-model-name Ornith-1.0-35B \
    --tensor-parallel-size 8 \
    --host 0.0.0.0 --port 8000 \
    --max-model-len 262144 \
    --gpu-memory-utilization 0.90 \
    --enable-prefix-caching \
    --enable-auto-tool-choice --tool-call-parser qwen3_xml \
    --reasoning-parser qwen3 \
    --trust-remote-code

SGLang

python -m sglang.launch_server \
    --model-path deepreinforce-ai/Ornith-1.0-35B \
    --served-model-name Ornith-1.0-35B \
    --tp 8 \
    --host 0.0.0.0 --port 8000 \
    --context-length 262144 \
    --mem-fraction-static 0.85 \
    --tool-call-parser qwen3_coder \
    --reasoning-parser qwen3

Hugging Face Transformers

Requires transformers >= 5.8.1. See the Transformers installation guide.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "deepreinforce-ai/Ornith-1.0-35B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto", device_map="auto")

Agentic Coding Benchmark Results

Agentic Coding Benchmark Results

Benchmark Ornith-1.0-35B Qwen3.5-35B Qwen3.6-35B Gemma4-31B Qwen3.5-397B
Agentic Coding
Terminal-Bench 2.1 (Terminus-2) 64.2 41.4 52.5 42.1 53.5
Terminal-Bench 2.1 (Claude Code) 62.8 38.9 49.2 - 48.6
SWE-bench Verified 75.6 70 73.4 52 76.4
SWE-bench Pro 50.4 44.6 49.5 35.7 51.6
SWE-bench Multilingual 69.3 60.3 67.2 51.7 69.3
NL2Repo 34.6 20.5 29.4 15.5 36.8
Claw-eval Avg 69.8 65.4 68.7 48.5 70.7
SWE Atlas - QnA 37.1 13.2 15.5 - 20.4
SWE Atlas - RF 29.7 10.2 11.4 - 18.4
SWE Atlas - TW 27.8 9.8 13.3 - 18.5

Evaluation Methodology

  • Terminal-Bench 2.1 (Terminus-2): Harbor/Terminus-2 framework, parser=json, temperature=1.0, top_p=1.0, 128K context, 4h timeout, 32 CPU cores, 48GB RAM, avg of 5 runs.
  • Terminal-Bench 2.1 (Claude Code): Claude Code 2.1.126, parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072, avg of 5 runs.
  • SWE-Bench Verified/Pro/Multilingual: OpenHands harness, temp=1.0, top_p=0.95, 256k context window.
  • SWE Atlas QnA/RF/TW: mini SWE agent harness, temp=1.0, top_p=0.95, 128K context, avg of 5 runs.
  • NL2Repo: temperature=1.0, top_p=1.0, 400K context, 48K output, anti-hacking filters.
  • ClawEval: Agentic code benchmark over real-user task distributions, temp=0.6, 256K context.

Key Highlights

Ornith-1.0 Key Highlights

Ornith-1.0 is a self-improving family of open-source models for agentic coding.

  • State-of-the-Art Coding Agents: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
  • Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scaffold that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.
  • Licence: MIT licensed, globally accessible, and free from regional limitations.

The 35B-MoE variant is the lightweight member of the Ornith family, designed for efficient single-GPU deployment.

Architecture

Decoder Block ×40 input Embedding vocab 248K · d 2048 Full Attention Hybrid 16:2 · dₕ 256 ×10 Linear / Recurrent Hybrid 16:2 · dₕ 256 ×30 MoE FFN 256 experts · top-8 · +1 shared · dᴻ 512 MTP Head ×1 speculative layer Final RMSNorm LM Head vocab 248K output
Attention
Hybrid Attention (16:2)
MoE
256 experts · top-8 per token
Layers
40
Hidden size
2048
Context
262K tokens
RoPE θ
10M
Parameters
35000M
Active params
3000M

Source: extracted model record · model repo

Type: Sparse MoE hybrid linear-attention decoder (Ornith-1.0-35B, Qwen3.5 base)
Attention: Hybrid: 30 linear-attention layers (Gated DeltaNet style, conv kernel 4, key heads 16x128, value heads 32x128) + 10 full-attention layers (interval 4; 16 Q / 2 KV heads, head dim 256, partial RoPE 0.25, theta 10M, interleaved MRoPE [11,11,10])
Decoder: Sparse MoE decoder-only with linear-attention/full hybrid (Qwen3_5Moe)
MoE: yes (256 experts)
Routing: Top-8 of 256 routed experts + 1 shared expert (shared expert FFN 512)
Layers 40
Total parameters 35000M
Active parameters 3000M
Context length 262K
Extended context 262K
Experts 256
Experts per token 8
Shared experts 1
Attention heads 16
KV heads 2
Head dim 256
Hidden size 2048
Vocabulary 248K
Expert FFN dim 512
Precision bfloat16
MTP layers 1
Vision encoder qwen3_5_moe_vision ViT (depth 27, hidden 1152, patch 16, spatial merge 2, out 2048)
RoPE θ 10M
Full Attention Interval 4
Full Attention Layers 10
Linear Attention Layers 30
Linear Conv Kernel Dim 4
Linear Key Head Dim 128
Linear Num Key Heads 16
Linear Num Value Heads 32
Linear Value Head Dim 128
Modalities text+image in, text out
Model type qwen3_5_moe
Mrope Section 11, 11, 10
Partial Rotary Factor 0.25
Shared Expert Intermediate Size 512

Training Pipeline

  1. 1
    cpt

    Continued pretraining on top of Qwen3.5 and Gemma 4

    Additional continued pretraining, mid-training and post-training on top of Gemma 4 and Qwen 3.5 bases (DB training_data_info).

  2. 2
    other

    Mid-training

    Mid-training stage between CPT and post-training (DB training_data_info).

  3. 3
    rl

    Self-improving scaffold-rollout RL (post-training)

    RL jointly optimizing scaffolds and solution rollouts; discovers better search trajectories, higher-quality solutions (card Self-Improving Training Framework).

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Agentic coding RL rollouts (solution + scaffold) rl — —

Trend Analysis

24h Change

+7.6%

7d Change

+665.0%

Current

282,300 pulls

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 282,300 pulls daily 01.09.2026
ollama downloads 262,300 pulls daily 31.08.2026
ollama downloads 198,400 pulls daily 30.08.2026
ollama downloads 142,200 pulls daily 29.08.2026
ollama downloads 117,900 pulls daily 28.08.2026
ollama downloads 65,100 pulls daily 27.08.2026
ollama downloads 44,800 pulls daily 26.08.2026
ollama downloads 36,900 pulls daily 25.08.2026
ollama downloads 31,000 pulls daily 24.08.2026

View full metric history →

Related Models