Parameters
36.0B total / 3.0B active
MoE: total / active
Architecture
Mixture-of-Experts (qwen3_5_moe)
Released
18.08.2026
License
MIT License
Input Modalities
Output Modalities
Context (native)
262,144 tokens
Context (extended)
1,048,576 tokens
About
Ornith-1.5-35B-A3B (ornith-ai/Ornith-1.5-35B-A3B) is the mid-size mixture-of-experts member of the Ornith-1.5 family - ~36B total parameters with only ~3B activated per token (qwen3_5_moe architecture, built on Qwen3.5 and Gemma4 with continued pretraining, mid-training and post-training) - released August 18, 2026 under the MIT License with a 262,144-token native context extensible to 1,048,576.
Ornith-1.5 is a major step toward foundation models through end-to-end self-improvement: it extends Ornith-1.0's scaffold-rollout co-optimization (jointly optimizing scaffold and solution rollouts) by also optimizing task generation. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. Despite activating only ~3B parameters per token, it significantly outperforms its similar-sized peer Qwen 3.6-35B across all coding and agentic benchmarks, and beats dense models such as Gemma 4-31B and Muse Glimmer-30B by wide margins on agentic coding. GGUF builds are available for llama.cpp, Ollama, Atomic.chat, Hermes, OpenClaw and coding CLIs.
Training Data Self-improvement loop jointly optimizing task generation, scaffold construction, and solution rollouts. Built on Qwen3.5 and Gemma4 with continued pretraining, mid-training, and post-training. Continuously generates new training tasks, discovers effective strategies, and improves policy through reinforcement learning.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
69.03%
|
19.08.2026 |
|
Terminal-Bench 2.1 (Claude Code)
coding_agent
|
100.00%
|
19.08.2026 |
|
SWE-bench Verified
coding_agent
|
90.14%
|
19.08.2026 |
|
SWE-bench Pro
coding_agent
|
74.50%
|
19.08.2026 |
|
SWE-bench Multilingual
coding_agent
|
78.58%
|
19.08.2026 |
|
DeepSWE
coding_agent
|
30.26%
|
19.08.2026 |
|
Frontier-Bench v0.1
coding_agent
|
100.00%
|
19.08.2026 |
|
NL2Repo
coding_agent
|
55.23%
|
19.08.2026 |
|
SWE Atlas - QnA
coding_agent
|
55.84%
|
19.08.2026 |
|
Humanity's Last Exam
stem_reasoning
|
44.51%
|
19.08.2026 |
|
HLE (with tools)
stem_reasoning
|
34.11%
|
19.08.2026 |
|
GPQA Diamond
stem_reasoning
|
90.36%
|
19.08.2026 |
|
MCP-Atlas
general_agent
|
78.80%
|
19.08.2026 |
|
Toolathlon Verified
general_agent
|
41.72%
|
19.08.2026 |
|
WideSearch
general_agent
|
63.80%
|
19.08.2026 |
|
BrowseComp
general_agent
|
73.21%
|
19.08.2026 |
|
WildClawBench
coding_agent
|
100.00%
|
19.08.2026 |
Ornith-1.5-35B-A3B: Thinking Control & Reasoning Mode
Thinking Control & Reasoning Mode — Ornith-1.5-35B-A3B
Ornith-1.5-35B-A3B is a reasoning model: by default the assistant turn opens with a <think>...</think> block before the final answer.
Reasoning Parser
The serving recipes enable a reasoning parser (qwen3) so the chain-of-thought is returned in a separate reasoning_content field:
- vLLM:
--reasoning-parser qwen3 - SGLang:
--reasoning-parser qwen3
Tool-Call Parser
A tool-call parser surfaces the model's <tool_call> blocks as OpenAI-style tool_calls:
- vLLM:
--tool-call-parser qwen3_xml - SGLang:
--tool-call-parser qwen3_coder
Chat Template
The model uses a modified Qwen chat template to ensure consistency between training and inference. The custom template is available at: https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B/blob/main/chat_template.jinja
Ornith-1.5-35B-A3B: Model Ecosystem (Finetunes, Merges, Quantizations)
Model Ecosystem — Ornith-1.5-35B-A3B
Model Tree
Ornith-1.5-35B-A3B has the following community-derived variants on HuggingFace:
| Type | Count |
|---|---|
| Finetunes | 13 models |
| Merges | 1 model |
| Quantizations | 82 models |
Quantization Variants
Notable official quantization repositories:
- FP8: ornith-ai/Ornith-1.5-35B-A3B-FP8
- GGUF: ornith-ai/Ornith-1.5-35B-A3B-GGUF (BF16 gguf: 71.1 GB, mmproj: 903 MB)
- NVFP4: ornith-ai/Ornith-1.5-35B-A3B-NVFP4
Collection
Part of the Ornith-1.5 collection (12 items, updated 5 days ago, 115 total).
Ornith-1.5-35B-A3B: Citation
Citation — Ornith-1.5-35B-A3B
@misc{ornith_1_5,
title = {{Ornith-1.5}: From Self-Scaffolding to Self-Improvement},
url = {https://ornith.ai/ornith_1_5.html},
author = {{Ornith Team}},
year = {2026}
}
Blog post: https://ornith.ai/ornith_1_5.html
Ornith-1.5-35B-A3B: Recommended Sampling Parameters & Best Practices
Recommended Sampling Parameters & Best Practices — Ornith-1.5-35B-A3B
Sampling Parameters
| Use Case | Temperature | top_p | top_k |
|---|---|---|---|
| General tasks | 0.6 | 0.95 | 20 |
| Reproduce reported benchmarks | 1.0 | (as per benchmark spec) | (as per benchmark spec) |
Key Guidelines
- Reasoning model: By default the assistant turn opens with a
<think>...</think>block before the final answer. The serving recipes enable a reasoning parser so the chain-of-thought is returned in a separatereasoning_contentfield. - Tool-call parser: Use
qwen3_xml(vLLM) orqwen3_coder(SGLang) to surface the model's<tool_call>blocks as OpenAI-styletool_calls. - Reasoning parser: Use
qwen3for both vLLM and SGLang. - Context window: 262,144 tokens native. Only enable YaRN RoPE scaling when your workload genuinely needs longer context.
- YaRN factor sizing: Target window ≈
factor× 262,144. For 524K tokens, usefactor: 2.0; for 1M tokens, usefactor: 4.0. - GPU memory: 2× 80GB GPUs recommended for 256K context with
gpu-memory-utilization 0.90(vLLM) ormem-fraction-static 0.85(SGLang).
Ornith-1.5-35B-A3B: Context Length & YaRN Long-Context Extension
Context Length & Long-Context Extension — Ornith-1.5-35B-A3B
Native Context
Ornith-1.5-35B-A3B handles context windows of up to 262,144 tokens natively.
Extended Context via YaRN RoPE Scaling
When a task's combined input and output must go beyond the native limit, the effective window can be extended with RoPE scaling — YaRN is the validated technique, already built into both vLLM and SGLang.
With a scaling factor of 4.0, the usable window grows to roughly 1M tokens.
Method 1: Edit checkpoint's config.json
{
"rope_scaling": {
"rope_type": "yarn",
"factor": 4.0,
"original_max_position_embeddings": 262144
}
}
Method 2: Override at launch time
vLLM:
VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ornith-ai/Ornith-1.5-35B-A3B ... \
--hf-overrides '{"rope_scaling": {"rope_type": "yarn", "factor": 4.0, "original_max_position_embeddings": 262144}}' \
--max-model-len 1000000
SGLang:
SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 python -m sglang.launch_server ... \
--json-model-override-args '{"rope_scaling": {"rope_type": "yarn", "factor": 4.0, "original_max_position_embeddings": 262144}}' \
--context-length 1000000
Note: Open-source runtimes implement YaRN statically: the same scaling factor is applied to every request regardless of its length, which can slightly hurt quality on ordinary-length inputs. Only enable
rope_scalingwhen your workload genuinely needs the longer window, and sizefactorto match it — the target window is roughlyfactor× 262,144. For requests topping out around 524,288 tokens,factor: 2.0is the better setting.
Ornith-1.5-35B-A3B: Chat Completions API Usage with Tool Calling
Chat Completions API Usage — Ornith-1.5-35B-A3B
Once a vLLM or SGLang server is running, talk to it with any OpenAI-compatible client.
Basic Usage
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="Ornith-1.5-35B-A3B",
messages=[
{"role": "user", "content": "Write a one-line Python lambda that squares a number."}
],
temperature=0.6,
top_p=0.95,
max_tokens=1024,
)
message = response.choices[0].message
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)
Tool Calling
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
]
response = client.chat.completions.create(
model="Ornith-1.5-35B-A3B",
messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
tools=tools,
tool_choice="auto",
temperature=0.6,
max_tokens=2048,
)
tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
## -> get_weather {"city": "Paris"}
The model emits well-formed function calls that the server parses into the standard tool_calls field. Streaming is also supported. Any OpenAI-compatible SDK (Python, Node.js, etc.) or curl can be pointed at the same /v1/chat/completions endpoint.
Ornith-1.5-35B-A3B: Deployment & Serving Recipes (vLLM, SGLang, Ollama, llama.cpp)
Deployment & Serving Recipes — Ornith-1.5-35B-A3B
Ornith-1.5-35B-A3B is a ~35B mixture-of-experts model with ~3B activated parameters per token (≈70 GB in bf16). The recipes below stand up an OpenAI-compatible server on 2× 80GB GPUs to leave headroom for the 256K context.
Runtime Requirements
- Transformers ≥ 5.8.1
- vLLM ≥ 0.19.1
- SGLang ≥ 0.5.9
vLLM
vllm serve ornith-ai/Ornith-1.5-35B-A3B \
--served-model-name Ornith-1.5-35B-A3B \
--host 0.0.0.0 --port 8000 \
--tensor-parallel-size 2 \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--trust-remote-code
SGLang
python -m sglang.launch_server \
--model-path ornith-ai/Ornith-1.5-35B-A3B \
--served-model-name Ornith-1.5-35B-A3B \
--host 0.0.0.0 --port 8000 \
--tp 2 \
--context-length 262144 \
--mem-fraction-static 0.85 \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3
Ollama
ollama run hf.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF
llama.cpp
llama-server -hf ornith-ai/Ornith-1.5-35B-A3B-GGUF --port 8000 -c 262144
Unsloth Studio
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
"ornith-ai/Ornith-1.5-35B-A3B",
max_seq_length=262144,
load_in_4bit=True,
)
Agent Framework Integration
Works with Hermes Agent, OpenClaw, Atomic.chat, and any OpenAI-compatible agent framework. Point the framework at your Ornith server via OPENAI_BASE_URL and OPENAI_API_KEY.
Ornith-1.5-35B-A3B: Benchmark Results Across Coding, Reasoning, and Agentic Tasks
Benchmark Results — Ornith-1.5-35B-A3B
All results are averaged over five independent runs.
Coding
| Benchmark | Ornith-1.5-35B-A3B | Ornith-1.0-35B-A3B | Qwen3.6-35B-A3B | Gemma-4-31B | Muse-Glimmer-30B | Qwen3.5-397B |
|---|---|---|---|---|---|---|
| Terminal-Bench 2.1 (Terminus-2) | 67.8 | 64.2 | 52.5 | 42.1 | 51.7 | 53.5 |
| Terminal-Bench 2.1 (Claude Code) | 68.5 | 62.8 | 49.2 | - | - | 48.6 |
| SWE-bench Verified | 79 | 75.6 | 73.4 | 52 | 76 | 76.4 |
| SWE-bench Pro | 59.6 | 50.4 | 49.5 | 35.7 | 51.2 | 51.6 |
| SWE-bench Multilingual | 71.4 | 69.3 | 67.2 | 51.7 | - | 69.3 |
| DeepSWE | 22 | 0 | 0 | - | - | 1 |
| Frontier-Bench v0.1 | 5.1 | 1.4 | 1.4 | - | - | 1.4 |
| NL2Repo | 46.2 | 34.6 | 29.4 | 15.5 | - | 36.8 |
| SWE Atlas - QnA | 39.8 | 37.1 | 15.5 | - | - | 20.4 |
Reasoning
| Benchmark | Ornith-1.5-35B-A3B | Ornith-1.0-35B-A3B | Qwen3.6-35B-A3B | Gemma-4-31B | Muse-Glimmer-30B | Qwen3.5-397B |
|---|---|---|---|---|---|---|
| HLE (no tools) | 25.6 | 20.8 | 21.4 | 19.5 | 22 | 28.7 |
| HLE (with tools) | 33.4 | 30.1 | 28.9 | 26.5 | - | 48.3 |
| GPQA Diamond | 89.2 | 86.2 | 86 | 84.3 | 83.5 | 88.4 |
Agentic
| Benchmark | Ornith-1.5-35B-A3B | Ornith-1.0-35B-A3B | Qwen3.6-35B-A3B | Gemma-4-31B | Muse-Glimmer-30B | Qwen3.5-397B |
|---|---|---|---|---|---|---|
| MCP-Atlas | 70.2 | 64.4 | 62.8 | 55 | 75.5 | 72.3 |
| Toolathlon-Verified | 48.7 | 42.4 | 41.7 | 40.8 | - | 38.3 |
| WideSearch | 67.8 | 63.4 | 60.1 | 54.2 | - | 74 |
| BrowseComp | 67.6 | 63.5 | 62 | - | - | 78.6 |
| ClawEval | 72.5 | 69.8 | 68.7 | 48.5 | - | 70.7 |
Evaluation Notes
- Terminal-Bench 2.1 (Terminus-2): Harbor/Terminus-2 framework, parser=json, temp=1.0, top_p=1.0, 128K context, 4-hour timeout, 32 CPU cores, 48GB RAM, 5 runs averaged.
- Terminal-Bench 2.1 (Claude Code): Claude Code 2.1.126, parser=json, temp=1.0, top_p=1.0, max_new_tokens=131072, 5 runs averaged.
- SWE-Bench Verified/Pro/Multilingual: OpenHands harness, temp=1.0, top_p=0.95, 256K context. Anti-hacking: Git history removed, network disabled.
- DeepSWE: Claude Code harness, temp=1.0, top_p=0.95, 256K context.
- SWE Atlas QnA: mini SWE agent harness, temp=1.0, top_p=0.95, 128K context, 5 runs averaged.
- NL2Repo: temp=1.0, top_p=1.0, 400K context, 48K output. GitHub repos and pip packages blocked.
- HLE: Claude 4.6 Opus as judge model.
- MCP-Atlas: 500-task public subset, thinking mode, 10-min timeout, Claude 4.8 Opus as judge.
- Toolathlon-Verified: Official evaluation service, 128K max tokens.
- ClawEval: Agentic code benchmark, temp=0.6, 256K context.
Ornith-1.5-35B-A3B: Self-Improvement Loop & Key Highlights
Ornith-1.5-35B-A3B — Overview & Highlights
Ornith-1.5 is a major step toward building foundation models through end-to-end self-improvement. It extends Ornith-1.0 (developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts.
Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning.
Key claims
- Mid-size MoE member of the Ornith-1.5 family: ~35B total params, ~3B activated per token
- Significantly outperforms Qwen 3.6-35B across all coding and agentic benchmarks
- Outperforms dense models Gemma 4-31B and Muse Glimmer-30B by wide margins on agentic coding
- Self-improvement loop: jointly optimizes task generation, scaffold construction, and solution rollouts
Blog: https://ornith.ai/ornith_1_5.html
Architecture
- Attention
- Hybrid Attention (16:2)
- MoE
- 256 experts · top-8 per token
- Layers
- 40
- Hidden size
- 2048
- Context
- 262K tokens
- Parameters
- 36000M
- Active params
- 3000M
Source: Hugging Face config.json · Qwen3_5MoeForConditionalGeneration · exact layer pattern · model repo
qwen3_5_moe
Training Pipeline
-
1
cpt
Continued Pretraining on Qwen3.5/Gemma4 base
Ornith-1.0 was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training.
-
2
sft
Mid-training and SFT
Mid-training phase with scaffold and rollout optimization to build foundational agentic capabilities.
-
3
rl
Self-improvement RL loop (Ornith-1.5)
Ornith-1.5 expands the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| Self-generated RL tasks (task generation + scaffold + solution rollouts) | rl | — | — |
Linked Resources
Ornith-1.5 Blog
https://ornith.ai/ornith_1_5.html
Ornith Blog (Deep Reinforce)
https://deep-reinforce.com/ornith.html
Ornith-1.5-35B-A3B HuggingFace Model Card
https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B
Ornith-1.5-35B-A3B-GGUF (Quantized)
https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF
GPQA Diamond Leaderboard
https://huggingface.co/datasets/Idavidrein/gpqa
SWE-bench Verified Leaderboard
https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified
SWE-bench Pro Leaderboard
https://huggingface.co/datasets/ScaleAI/SWE-bench_Pro
DeepSWE Leaderboard
https://huggingface.co/datasets/datacurve/deep-swe
SWE-bench Multilingual Leaderboard
https://huggingface.co/datasets/SWE-bench/SWE-bench_Multilingual
HLE Leaderboard
https://huggingface.co/datasets/cais/hle
Trend Analysis
24h Change
+0.9%
7d Change
+9.3%
Current
3,185
likes
+1.9%
downloads
+19.7%
downloads_all_time
+19.7%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| huggingface | downloads_all_time | 206,651 | daily | 01.09.2026 |
| huggingface | followers | 3,185 | daily | 01.09.2026 |
| huggingface | likes | 523 | daily | 01.09.2026 |
| huggingface | downloads | 206,651 | daily | 01.09.2026 |
| huggingface | downloads_all_time | 172,695 | daily | 31.08.2026 |
| huggingface | followers | 3,158 | daily | 31.08.2026 |
| huggingface | likes | 513 | daily | 31.08.2026 |
| huggingface | downloads | 172,695 | daily | 31.08.2026 |
| huggingface | downloads_all_time | 147,038 | daily | 30.08.2026 |
| huggingface | followers | 3,129 | daily | 30.08.2026 |
| huggingface | likes | 505 | daily | 30.08.2026 |
| huggingface | downloads | 147,038 | daily | 30.08.2026 |
| huggingface | followers | 3,096 | daily | 29.08.2026 |
| huggingface | likes | 495 | daily | 29.08.2026 |
| huggingface | downloads | 106,562 | daily | 29.08.2026 |
| huggingface | followers | 3,061 | daily | 28.08.2026 |
| huggingface | likes | 483 | daily | 28.08.2026 |
| huggingface | downloads | 88,102 | daily | 28.08.2026 |
| huggingface | followers | 3,020 | daily | 27.08.2026 |
| huggingface | likes | 464 | daily | 27.08.2026 |
| huggingface | downloads | 88,102 | daily | 27.08.2026 |
| huggingface | followers | 2,976 | daily | 26.08.2026 |
| huggingface | likes | 451 | daily | 26.08.2026 |
| huggingface | downloads | 83,342 | daily | 26.08.2026 |
| huggingface | followers | 2,913 | daily | 25.08.2026 |
| huggingface | likes | 418 | daily | 25.08.2026 |
| huggingface | downloads | 70,158 | daily | 25.08.2026 |
| huggingface | followers | 2,860 | daily | 24.08.2026 |
| huggingface | likes | 394 | daily | 24.08.2026 |
| huggingface | downloads | 60,294 | daily | 24.08.2026 |
| huggingface | followers | 2,803 | daily | 23.08.2026 |
| huggingface | likes | 361 | daily | 23.08.2026 |
| huggingface | downloads | 23,516 | daily | 23.08.2026 |
| huggingface | followers | 2,741 | daily | 22.08.2026 |
| huggingface | likes | 320 | daily | 22.08.2026 |
| huggingface | downloads | 12,611 | daily | 22.08.2026 |