Ornith-1.5-35B-A3B

Ornith AI

Parameters

36.0B total / 3.0B active

MoE: total / active

Architecture

Mixture-of-Experts (qwen3_5_moe)

Released

18.08.2026

License

MIT License

Open Weights Commercial Use Multimodal BF16 Ornith-1.5 en multi

Input Modalities

text image

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

1,048,576 tokens

Openness Index Score 100.0/100

About

Ornith-1.5-35B-A3B (ornith-ai/Ornith-1.5-35B-A3B) is the mid-size mixture-of-experts member of the Ornith-1.5 family - ~36B total parameters with only ~3B activated per token (qwen3_5_moe architecture, built on Qwen3.5 and Gemma4 with continued pretraining, mid-training and post-training) - released August 18, 2026 under the MIT License with a 262,144-token native context extensible to 1,048,576.

Ornith-1.5 is a major step toward foundation models through end-to-end self-improvement: it extends Ornith-1.0's scaffold-rollout co-optimization (jointly optimizing scaffold and solution rollouts) by also optimizing task generation. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. Despite activating only ~3B parameters per token, it significantly outperforms its similar-sized peer Qwen 3.6-35B across all coding and agentic benchmarks, and beats dense models such as Gemma 4-31B and Muse Glimmer-30B by wide margins on agentic coding. GGUF builds are available for llama.cpp, Ollama, Atomic.chat, Hermes, OpenClaw and coding CLIs.

Training Data Self-improvement loop jointly optimizing task generation, scaffold construction, and solution rollouts. Built on Qwen3.5 and Gemma4 with continued pretraining, mid-training, and post-training. Continuously generates new training tasks, discovers effective strategies, and improves policy through reinforcement learning.

Benchmark Scores

Benchmark Score Date
Terminal-Bench 2.1 (Terminus-2)
coding_agent
69.03%
19.08.2026
Terminal-Bench 2.1 (Claude Code)
coding_agent
100.00%
19.08.2026
SWE-bench Verified
coding_agent
90.14%
19.08.2026
SWE-bench Pro
coding_agent
74.50%
19.08.2026
SWE-bench Multilingual
coding_agent
78.58%
19.08.2026
DeepSWE
coding_agent
30.26%
19.08.2026
Frontier-Bench v0.1
coding_agent
100.00%
19.08.2026
NL2Repo
coding_agent
55.23%
19.08.2026
SWE Atlas - QnA
coding_agent
55.84%
19.08.2026
Humanity's Last Exam
stem_reasoning
44.51%
19.08.2026
HLE (with tools)
stem_reasoning
34.11%
19.08.2026
GPQA Diamond
stem_reasoning
90.36%
19.08.2026
MCP-Atlas
general_agent
78.80%
19.08.2026
Toolathlon Verified
general_agent
41.72%
19.08.2026
WideSearch
general_agent
63.80%
19.08.2026
BrowseComp
general_agent
73.21%
19.08.2026
WildClawBench
coding_agent
100.00%
19.08.2026

Ornith-1.5-35B-A3B: Thinking Control & Reasoning Mode

Thinking Control & Reasoning Mode — Ornith-1.5-35B-A3B

Ornith-1.5-35B-A3B is a reasoning model: by default the assistant turn opens with a <think>...</think> block before the final answer.

Reasoning Parser

The serving recipes enable a reasoning parser (qwen3) so the chain-of-thought is returned in a separate reasoning_content field:

  • vLLM: --reasoning-parser qwen3
  • SGLang: --reasoning-parser qwen3

Tool-Call Parser

A tool-call parser surfaces the model's <tool_call> blocks as OpenAI-style tool_calls:

  • vLLM: --tool-call-parser qwen3_xml
  • SGLang: --tool-call-parser qwen3_coder

Chat Template

The model uses a modified Qwen chat template to ensure consistency between training and inference. The custom template is available at: https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B/blob/main/chat_template.jinja

Ornith-1.5-35B-A3B: Model Ecosystem (Finetunes, Merges, Quantizations)

Model Ecosystem — Ornith-1.5-35B-A3B

Model Tree

Ornith-1.5-35B-A3B has the following community-derived variants on HuggingFace:

Type Count
Finetunes 13 models
Merges 1 model
Quantizations 82 models

Quantization Variants

Notable official quantization repositories:

Collection

Part of the Ornith-1.5 collection (12 items, updated 5 days ago, 115 total).

Ornith-1.5-35B-A3B: Citation

Citation — Ornith-1.5-35B-A3B

@misc{ornith_1_5,
    title = {{Ornith-1.5}: From Self-Scaffolding to Self-Improvement},
    url = {https://ornith.ai/ornith_1_5.html},
    author = {{Ornith Team}},
    year = {2026}
}

Blog post: https://ornith.ai/ornith_1_5.html

Ornith-1.5-35B-A3B: Recommended Sampling Parameters & Best Practices

Recommended Sampling Parameters & Best Practices — Ornith-1.5-35B-A3B

Sampling Parameters

Use Case Temperature top_p top_k
General tasks 0.6 0.95 20
Reproduce reported benchmarks 1.0 (as per benchmark spec) (as per benchmark spec)

Key Guidelines

  • Reasoning model: By default the assistant turn opens with a <think>...</think> block before the final answer. The serving recipes enable a reasoning parser so the chain-of-thought is returned in a separate reasoning_content field.
  • Tool-call parser: Use qwen3_xml (vLLM) or qwen3_coder (SGLang) to surface the model's <tool_call> blocks as OpenAI-style tool_calls.
  • Reasoning parser: Use qwen3 for both vLLM and SGLang.
  • Context window: 262,144 tokens native. Only enable YaRN RoPE scaling when your workload genuinely needs longer context.
  • YaRN factor sizing: Target window ≈ factor × 262,144. For 524K tokens, use factor: 2.0; for 1M tokens, use factor: 4.0.
  • GPU memory: 2× 80GB GPUs recommended for 256K context with gpu-memory-utilization 0.90 (vLLM) or mem-fraction-static 0.85 (SGLang).

Ornith-1.5-35B-A3B: Context Length & YaRN Long-Context Extension

Context Length & Long-Context Extension — Ornith-1.5-35B-A3B

Native Context

Ornith-1.5-35B-A3B handles context windows of up to 262,144 tokens natively.

Extended Context via YaRN RoPE Scaling

When a task's combined input and output must go beyond the native limit, the effective window can be extended with RoPE scaling — YaRN is the validated technique, already built into both vLLM and SGLang.

With a scaling factor of 4.0, the usable window grows to roughly 1M tokens.

Method 1: Edit checkpoint's config.json

{
    "rope_scaling": {
        "rope_type": "yarn",
        "factor": 4.0,
        "original_max_position_embeddings": 262144
    }
}

Method 2: Override at launch time

vLLM:

VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ornith-ai/Ornith-1.5-35B-A3B ... \
    --hf-overrides '{"rope_scaling": {"rope_type": "yarn", "factor": 4.0, "original_max_position_embeddings": 262144}}' \
    --max-model-len 1000000

SGLang:

SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 python -m sglang.launch_server ... \
    --json-model-override-args '{"rope_scaling": {"rope_type": "yarn", "factor": 4.0, "original_max_position_embeddings": 262144}}' \
    --context-length 1000000

Note: Open-source runtimes implement YaRN statically: the same scaling factor is applied to every request regardless of its length, which can slightly hurt quality on ordinary-length inputs. Only enable rope_scaling when your workload genuinely needs the longer window, and size factor to match it — the target window is roughly factor × 262,144. For requests topping out around 524,288 tokens, factor: 2.0 is the better setting.

Ornith-1.5-35B-A3B: Chat Completions API Usage with Tool Calling

Chat Completions API Usage — Ornith-1.5-35B-A3B

Once a vLLM or SGLang server is running, talk to it with any OpenAI-compatible client.

Basic Usage

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="EMPTY",
)

response = client.chat.completions.create(
    model="Ornith-1.5-35B-A3B",
    messages=[
        {"role": "user", "content": "Write a one-line Python lambda that squares a number."}
    ],
    temperature=0.6,
    top_p=0.95,
    max_tokens=1024,
)

message = response.choices[0].message
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)

Tool Calling

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a city",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
                "required": ["city"],
            },
        },
    }
]

response = client.chat.completions.create(
    model="Ornith-1.5-35B-A3B",
    messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
    tools=tools,
    tool_choice="auto",
    temperature=0.6,
    max_tokens=2048,
)

tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
## -> get_weather {"city": "Paris"}

The model emits well-formed function calls that the server parses into the standard tool_calls field. Streaming is also supported. Any OpenAI-compatible SDK (Python, Node.js, etc.) or curl can be pointed at the same /v1/chat/completions endpoint.

Ornith-1.5-35B-A3B: Deployment & Serving Recipes (vLLM, SGLang, Ollama, llama.cpp)

Deployment & Serving Recipes — Ornith-1.5-35B-A3B

Ornith-1.5-35B-A3B is a ~35B mixture-of-experts model with ~3B activated parameters per token (≈70 GB in bf16). The recipes below stand up an OpenAI-compatible server on 2× 80GB GPUs to leave headroom for the 256K context.

Runtime Requirements

  • Transformers ≥ 5.8.1
  • vLLM ≥ 0.19.1
  • SGLang ≥ 0.5.9

vLLM

vllm serve ornith-ai/Ornith-1.5-35B-A3B \
    --served-model-name Ornith-1.5-35B-A3B \
    --host 0.0.0.0 --port 8000 \
    --tensor-parallel-size 2 \
    --max-model-len 262144 \
    --gpu-memory-utilization 0.90 \
    --enable-prefix-caching \
    --enable-auto-tool-choice --tool-call-parser qwen3_xml \
    --reasoning-parser qwen3 \
    --trust-remote-code

SGLang

python -m sglang.launch_server \
    --model-path ornith-ai/Ornith-1.5-35B-A3B \
    --served-model-name Ornith-1.5-35B-A3B \
    --host 0.0.0.0 --port 8000 \
    --tp 2 \
    --context-length 262144 \
    --mem-fraction-static 0.85 \
    --tool-call-parser qwen3_coder \
    --reasoning-parser qwen3

Ollama

ollama run hf.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF

llama.cpp

llama-server -hf ornith-ai/Ornith-1.5-35B-A3B-GGUF --port 8000 -c 262144

Unsloth Studio

from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
    "ornith-ai/Ornith-1.5-35B-A3B",
    max_seq_length=262144,
    load_in_4bit=True,
)

Agent Framework Integration

Works with Hermes Agent, OpenClaw, Atomic.chat, and any OpenAI-compatible agent framework. Point the framework at your Ornith server via OPENAI_BASE_URL and OPENAI_API_KEY.

Ornith-1.5-35B-A3B: Benchmark Results Across Coding, Reasoning, and Agentic Tasks

Benchmark Results — Ornith-1.5-35B-A3B

All results are averaged over five independent runs.

Coding

Benchmark Ornith-1.5-35B-A3B Ornith-1.0-35B-A3B Qwen3.6-35B-A3B Gemma-4-31B Muse-Glimmer-30B Qwen3.5-397B
Terminal-Bench 2.1 (Terminus-2) 67.8 64.2 52.5 42.1 51.7 53.5
Terminal-Bench 2.1 (Claude Code) 68.5 62.8 49.2 - - 48.6
SWE-bench Verified 79 75.6 73.4 52 76 76.4
SWE-bench Pro 59.6 50.4 49.5 35.7 51.2 51.6
SWE-bench Multilingual 71.4 69.3 67.2 51.7 - 69.3
DeepSWE 22 0 0 - - 1
Frontier-Bench v0.1 5.1 1.4 1.4 - - 1.4
NL2Repo 46.2 34.6 29.4 15.5 - 36.8
SWE Atlas - QnA 39.8 37.1 15.5 - - 20.4

Reasoning

Benchmark Ornith-1.5-35B-A3B Ornith-1.0-35B-A3B Qwen3.6-35B-A3B Gemma-4-31B Muse-Glimmer-30B Qwen3.5-397B
HLE (no tools) 25.6 20.8 21.4 19.5 22 28.7
HLE (with tools) 33.4 30.1 28.9 26.5 - 48.3
GPQA Diamond 89.2 86.2 86 84.3 83.5 88.4

Agentic

Benchmark Ornith-1.5-35B-A3B Ornith-1.0-35B-A3B Qwen3.6-35B-A3B Gemma-4-31B Muse-Glimmer-30B Qwen3.5-397B
MCP-Atlas 70.2 64.4 62.8 55 75.5 72.3
Toolathlon-Verified 48.7 42.4 41.7 40.8 - 38.3
WideSearch 67.8 63.4 60.1 54.2 - 74
BrowseComp 67.6 63.5 62 - - 78.6
ClawEval 72.5 69.8 68.7 48.5 - 70.7

Evaluation Notes

  • Terminal-Bench 2.1 (Terminus-2): Harbor/Terminus-2 framework, parser=json, temp=1.0, top_p=1.0, 128K context, 4-hour timeout, 32 CPU cores, 48GB RAM, 5 runs averaged.
  • Terminal-Bench 2.1 (Claude Code): Claude Code 2.1.126, parser=json, temp=1.0, top_p=1.0, max_new_tokens=131072, 5 runs averaged.
  • SWE-Bench Verified/Pro/Multilingual: OpenHands harness, temp=1.0, top_p=0.95, 256K context. Anti-hacking: Git history removed, network disabled.
  • DeepSWE: Claude Code harness, temp=1.0, top_p=0.95, 256K context.
  • SWE Atlas QnA: mini SWE agent harness, temp=1.0, top_p=0.95, 128K context, 5 runs averaged.
  • NL2Repo: temp=1.0, top_p=1.0, 400K context, 48K output. GitHub repos and pip packages blocked.
  • HLE: Claude 4.6 Opus as judge model.
  • MCP-Atlas: 500-task public subset, thinking mode, 10-min timeout, Claude 4.8 Opus as judge.
  • Toolathlon-Verified: Official evaluation service, 128K max tokens.
  • ClawEval: Agentic code benchmark, temp=0.6, 256K context.

Ornith-1.5-35B-A3B: Self-Improvement Loop & Key Highlights

Ornith-1.5-35B-A3B — Overview & Highlights

Ornith-1.5 is a major step toward building foundation models through end-to-end self-improvement. It extends Ornith-1.0 (developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts.

Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning.

Key claims

  • Mid-size MoE member of the Ornith-1.5 family: ~35B total params, ~3B activated per token
  • Significantly outperforms Qwen 3.6-35B across all coding and agentic benchmarks
  • Outperforms dense models Gemma 4-31B and Muse Glimmer-30B by wide margins on agentic coding
  • Self-improvement loop: jointly optimizes task generation, scaffold construction, and solution rollouts

Blog: https://ornith.ai/ornith_1_5.html

Architecture

Decoder Block ×40 input Embedding vocab 248K · d 2048 Linear / Recurrent Hybrid 16:2 · dₕ 256 ×30 Full Attention Hybrid 16:2 · dₕ 256 ×10 MoE FFN 256 experts · top-8 · dᴻ 512 Final Norm LM Head vocab 248K output
Attention
Hybrid Attention (16:2)
MoE
256 experts · top-8 per token
Layers
40
Hidden size
2048
Context
262K tokens
Parameters
36000M
Active params
3000M

Source: Hugging Face config.json · Qwen3_5MoeForConditionalGeneration · exact layer pattern · model repo

Type: Mixture-of-Experts
Attention: Grouped Query Attention
Decoder: Causal Transformer
MoE: yes (128 experts)
Routing: Top-K router
Layers 48
Context length 262K
Extended context 1M
Experts 128
Experts per token 8
Hidden size 2560
Vision Yes
MTP Yes
RoPE dim 128
Base Architecture

qwen3_5_moe

Training Pipeline

  1. 1
    cpt

    Continued Pretraining on Qwen3.5/Gemma4 base

    Ornith-1.0 was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training.

  2. 2
    sft

    Mid-training and SFT

    Mid-training phase with scaffold and rollout optimization to build foundational agentic capabilities.

  3. 3
    rl

    Self-improvement RL loop (Ornith-1.5)

    Ornith-1.5 expands the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Self-generated RL tasks (task generation + scaffold + solution rollouts) rl — —

Trend Analysis

24h Change

+0.9%

7d Change

+9.3%

Current

3,185

huggingface

likes

+1.9%

huggingface

downloads

+19.7%

huggingface

downloads_all_time

+19.7%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
huggingface downloads_all_time 206,651 daily 01.09.2026
huggingface followers 3,185 daily 01.09.2026
huggingface likes 523 daily 01.09.2026
huggingface downloads 206,651 daily 01.09.2026
huggingface downloads_all_time 172,695 daily 31.08.2026
huggingface followers 3,158 daily 31.08.2026
huggingface likes 513 daily 31.08.2026
huggingface downloads 172,695 daily 31.08.2026
huggingface downloads_all_time 147,038 daily 30.08.2026
huggingface followers 3,129 daily 30.08.2026
huggingface likes 505 daily 30.08.2026
huggingface downloads 147,038 daily 30.08.2026
huggingface followers 3,096 daily 29.08.2026
huggingface likes 495 daily 29.08.2026
huggingface downloads 106,562 daily 29.08.2026
huggingface followers 3,061 daily 28.08.2026
huggingface likes 483 daily 28.08.2026
huggingface downloads 88,102 daily 28.08.2026
huggingface followers 3,020 daily 27.08.2026
huggingface likes 464 daily 27.08.2026
huggingface downloads 88,102 daily 27.08.2026
huggingface followers 2,976 daily 26.08.2026
huggingface likes 451 daily 26.08.2026
huggingface downloads 83,342 daily 26.08.2026
huggingface followers 2,913 daily 25.08.2026
huggingface likes 418 daily 25.08.2026
huggingface downloads 70,158 daily 25.08.2026
huggingface followers 2,860 daily 24.08.2026
huggingface likes 394 daily 24.08.2026
huggingface downloads 60,294 daily 24.08.2026
huggingface followers 2,803 daily 23.08.2026
huggingface likes 361 daily 23.08.2026
huggingface downloads 23,516 daily 23.08.2026
huggingface followers 2,741 daily 22.08.2026
huggingface likes 320 daily 22.08.2026
huggingface downloads 12,611 daily 22.08.2026

View full metric history →

Related Models