Qwen3.6-35B-A3B vs Ornith-1.0-35B

New Comparison

Scores are fetched live from the current database on every visit — nothing is cached.

Benchmark Scores Comparison

Qwen3.6-35B-A3B Ornith-1.0-35B

Qwen3.6-35B-A3B

35.0B params

Score: 100.0

Ornith-1.0-35B

35.0B params

Score: 100.0

Coverage matrix

2 models × 73 benchmarks — hover a benchmark name for its category

Full Coverage Matrix
Shading = score relative to the best score among the selected models: Low Mid High Missing (—) The totals row shows how many of the 73 benchmarks each model has scores for.
Benchmark Qwen3.6-35B-A3B Qwen Team (Alibaba Cloud) Ornith-1.0-35B DeepReinforce AI
HMMT Feb 25 90.7 —
DynaMath 82.8 —
GPQA Diamond 86 —
MathVista (mini) 86.4 —
MMBench EN-DEV v1.1 92.8 —
LiveCodeBench v6 80.4 —
SWE-bench Pro 49.5 50.4
RealWorldQA 85.3 —
EmbSpatialBench 84.3 —
AIME 26 92.7 —
Toolathlon Verified 41.7 —
VITA-Bench 35.6 —
TAU3-Bench 67.2 —
MMStar 80.7 —
SWE Atlas - TW 13.3 27.8
WildClawBench 68.7 —
SWE Atlas - QnA 15.5 37.1
AI2D_TEST 92.7 —
ParseBench Text Formatting 58.3 —
NL2Repo 29.4 34.6
MMMU-Pro 75.3 —
QwenClawBench 52.6 —
Claw-Eval Avg 68.7 69.8
VideoMMMU 83.7 —
HMMT Nov 25 89.1 —
MMLU-Pro 85.2 —
SuperGPQA 64.7 —
HLE (with tools) 28.9 —
LVBench 71.4 —
CharXiv (RQ) 78 —
BrowseComp 62 —
MMLU-Redux 93.3 —
ParseBench Text Content 90.7 —
MCPMark 37 —
SWE-bench Multilingual 67.2 69.3
MLVU 86.2 —
SWE-bench Verified 73.4 75.6
IMOAnswerBench 78.9 —
Humanity's Last Exam 21.4 —
CountBench 96.1 —
ERQA 61.8 —
VideoMME (w/o sub.) 82.5 —
Terminal-Bench 2.0 51.5 —
DeepPlanning 25.9 —
QwenWebBench 1397 —
ODInW13 50.8 —
VlmsAreBlind 96.6 —
Terminal-Bench 2.1 (Terminus-2) 52.5 64.2
DeepSWE 0 —
OmniDocBench 1.5 89.9 —
HallusionBench 69.8 —
OCRBench 90 —
HMMT Feb 26 83.6 —
CC-OCR 81.9 —
RefCOCO (avg) 92 —
Artificial Analysis Intelligence Index 32 —
Tool Decathlon 26.9 —
Frontier-Bench v0.1 1.4 —
ParseBench Mean 44.1 —
ZEROBench_sub 34.4 —
C-Eval 90 —
WideSearch 60.1 —
SimpleVQA 58.9 —
VideoMME (w sub.) 86.6 —
SkillsBench Avg5 28.7 —
Claw-Eval Pass^3 50 —
MCP-Atlas 62.8 —
SWE Atlas - RF 11.4 29.7
V-Star 90.1 —
Terminal-Bench 2.1 (Claude Code) 49.2 62.8
RefSpatialBench 64.3 —
MMMU 81.7 —
MVBench 74.6 —
Covered 73/73 10/73

Fully comparable benchmarks

Scored by all 2 selected models (10)

Benchmark Category Qwen3.6-35B-A3B Ornith-1.0-35B
SWE-bench Pro coding_agent 49.5 50.4
SWE Atlas - TW coding_agent 13.3 27.8
SWE Atlas - QnA coding_agent 15.5 37.1
NL2Repo coding_agent 29.4 34.6
Claw-Eval Avg coding_agent 68.7 69.8
SWE-bench Multilingual coding_agent 67.2 69.3
SWE-bench Verified coding_agent 73.4 75.6
Terminal-Bench 2.1 (Terminus-2) coding_agent 52.5 64.2
SWE Atlas - RF coding_agent 11.4 29.7
Terminal-Bench 2.1 (Claude Code) coding_agent 49.2 62.8

Partial coverage benchmarks

Scored by at least one but not all selected models (63) — missing scores show as —

Benchmark Category Qwen3.6-35B-A3B Ornith-1.0-35B
HMMT Feb 25 1/2 models stem_reasoning 90.7 —
DynaMath 1/2 models vision_language 82.8 —
GPQA Diamond 1/2 models stem_reasoning 86 —
MathVista (mini) 1/2 models vision_language 86.4 —
MMBench EN-DEV v1.1 1/2 models vision_language 92.8 —
LiveCodeBench v6 1/2 models stem_reasoning 80.4 —
RealWorldQA 1/2 models vision_language 85.3 —
EmbSpatialBench 1/2 models spatial_intelligence 84.3 —
AIME 26 1/2 models stem_reasoning 92.7 —
Toolathlon Verified 1/2 models general_agent 41.7 —
VITA-Bench 1/2 models general_agent 35.6 —
TAU3-Bench 1/2 models general_agent 67.2 —
MMStar 1/2 models vision_language 80.7 —
WildClawBench 1/2 models coding_agent 68.7 —
AI2D_TEST 1/2 models document_understanding 92.7 —
ParseBench Text Formatting 1/2 models document_understanding 58.3 —
MMMU-Pro 1/2 models vision_language 75.3 —
QwenClawBench 1/2 models coding_agent 52.6 —
VideoMMMU 1/2 models video_understanding 83.7 —
HMMT Nov 25 1/2 models stem_reasoning 89.1 —
MMLU-Pro 1/2 models knowledge 85.2 —
SuperGPQA 1/2 models knowledge 64.7 —
HLE (with tools) 1/2 models stem_reasoning 28.9 —
LVBench 1/2 models video_understanding 71.4 —
CharXiv (RQ) 1/2 models document_understanding 78 —
BrowseComp 1/2 models general_agent 62 —
MMLU-Redux 1/2 models knowledge 93.3 —
ParseBench Text Content 1/2 models document_understanding 90.7 —
MCPMark 1/2 models general_agent 37 —
MLVU 1/2 models video_understanding 86.2 —
IMOAnswerBench 1/2 models stem_reasoning 78.9 —
Humanity's Last Exam 1/2 models stem_reasoning 21.4 —
CountBench 1/2 models spatial_intelligence 96.1 —
ERQA 1/2 models spatial_intelligence 61.8 —
VideoMME (w/o sub.) 1/2 models video_understanding 82.5 —
Terminal-Bench 2.0 1/2 models coding_agent 51.5 —
DeepPlanning 1/2 models general_agent 25.9 —
QwenWebBench 1/2 models coding_agent 1397 —
ODInW13 1/2 models spatial_intelligence 50.8 —
VlmsAreBlind 1/2 models vision_language 96.6 —
DeepSWE 1/2 models coding_agent 0 —
OmniDocBench 1.5 1/2 models document_understanding 89.9 —
HallusionBench 1/2 models vision_language 69.8 —
OCRBench 1/2 models document_understanding 90 —
HMMT Feb 26 1/2 models stem_reasoning 83.6 —
CC-OCR 1/2 models document_understanding 81.9 —
RefCOCO (avg) 1/2 models spatial_intelligence 92 —
Artificial Analysis Intelligence Index 1/2 models composite 32 —
Tool Decathlon 1/2 models general_agent 26.9 —
Frontier-Bench v0.1 1/2 models coding_agent 1.4 —
ParseBench Mean 1/2 models document_understanding 44.1 —
ZEROBench_sub 1/2 models vision_language 34.4 —
C-Eval 1/2 models knowledge 90 —
WideSearch 1/2 models general_agent 60.1 —
SimpleVQA 1/2 models vision_language 58.9 —
VideoMME (w sub.) 1/2 models video_understanding 86.6 —
SkillsBench Avg5 1/2 models coding_agent 28.7 —
Claw-Eval Pass^3 1/2 models coding_agent 50 —
MCP-Atlas 1/2 models general_agent 62.8 —
V-Star 1/2 models vision_language 90.1 —
RefSpatialBench 1/2 models spatial_intelligence 64.3 —
MMMU 1/2 models vision_language 81.7 —
MVBench 1/2 models video_understanding 74.6 —