LLM Benchmarks — Scores, Methodology & Live Rankings
What each benchmark measures, how models score, and the current rankings
SWE-bench Verified
Agentic Coding
SWE-bench Verified evaluates LLM-based agents on real-world GitHub issue resolution from popular ...
SWE-bench Multilingual
Agentic Coding
SWE-bench Multilingual extends SWE-bench to multiple programming languages beyond Python.
SWE-bench Pro
Agentic Coding
SWE-bench Pro - refined benchmark with corrected problematic tasks in the public set. Internal ag...
Terminal-Bench 2.0
Agentic Coding
Terminal-Bench 2.0 - agentic coding and terminal use benchmark. Harbor/Terminus-2 harness; 3h tim...
Claw-Eval Pass^3
Agentic Coding
Claw-Eval Pass^3 - third-pass pass rate on the Claw-Eval agentic coding benchmark.
SkillsBench Avg5
Agentic Coding
SkillsBench Avg5 - evaluated via OpenCode on 78 self-contained tasks (excluding API-dependent tas...
QwenClawBench
Agentic Coding
QwenClawBench - internal real-user-distribution Claw agent benchmark (open-sourcing soon); temp=0...
NL2Repo
Agentic Coding
NL2Repo - repository generation from natural language. Others evaluated via Claude Code (temp=1.0...
QwenWebBench
Agentic Coding
QwenWebBench - internal front-end code generation benchmark; bilingual (EN/CN), 7 categories (Web...
TAU3-Bench
General Agents
TAU3-Bench - agentic tool use benchmark. Uses official user model (gpt-5.2, low reasoning effort)...
VITA-Bench
General Agents
VITA-Bench - agentic benchmark. Avg subdomain scores; using claude-4-sonnet as judger.
MCPMark
General Agents
MCPMark - GitHub MCP v0.30.3; Playwright responses truncated at 32K tokens.
MMLU-Pro
Knowledge
MMLU-Pro - extended MMLU covering 14 categories with complex reasoning questions.
C-Eval
Knowledge
C-Eval - comprehensive Chinese evaluation benchmark across 52 subjects.
GPQA Diamond
STEM & Reasoning
GPQA Diamond - Google-Proof Q&A benchmark for graduate-level scientific reasoning.
Humanity's Last Exam
STEM & Reasoning
HLE - Humanity's Last Exam - extremely difficult expert-level questions across domains.
LiveCodeBench v6
STEM & Reasoning
LiveCodeBench v6 - competitive programming evaluation with latest problems.
HMMT Feb 25
STEM & Reasoning
HMMT Feb 25 - Harvard-MIT Mathematics Tournament February 2025.
HMMT Nov 25
STEM & Reasoning
HMMT Nov 25 - Harvard-MIT Mathematics Tournament November 2025.
HMMT Feb 26
STEM & Reasoning
HMMT Feb 26 - Harvard-MIT Mathematics Tournament February 2026.
IMOAnswerBench
STEM & Reasoning
IMOAnswerBench - International Mathematical Olympiad answer evaluation benchmark.
AIME 26
STEM & Reasoning
AIME 26 - American Invitational Mathematics Examination 2026 (I & II).
MMMU
Vision & Language
MMMU - massive multi-discipline multimodal understanding and reasoning benchmark for college-leve...
MMMU-Pro
Vision & Language
MMMU-Pro - enhanced MMMU with harder questions and refined evaluation.
MathVista (mini)
Vision & Language
MathVista (mini) - mathematical reasoning in visual contexts.
RealWorldQA
Vision & Language
RealWorldQA - visual question answering on real-world images.
MMBench EN-DEV v1.1
Vision & Language
MMBench-EN-DEV v1.1 - multimodal benchmark English dev set.
HallusionBench
Vision & Language
HallusionBench - benchmark for language and visual hallucination.
OmniDocBench 1.5
Document Understanding
OmniDocBench 1.5 - document understanding and parsing benchmark.
CharXiv (RQ)
Document Understanding
CharXiv (RQ) - reasoning questions over arXiv figures and tables.
CC-OCR
Document Understanding
CC-OCR - comprehensive OCR benchmark across diverse content.
AI2D_TEST
Document Understanding
AI2D_TEST - grade-school science diagram understanding benchmark.
RefCOCO (avg)
Spatial Intelligence
RefCOCO (avg) - referring expression comprehension benchmark, average score.
ODInW13
Spatial Intelligence
ODInW13 - object detection in the wild across 13 domains.
EmbSpatialBench
Spatial Intelligence
EmbSpatialBench - embodied spatial intelligence benchmark.
RefSpatialBench
Spatial Intelligence
RefSpatialBench - referring spatial reasoning benchmark.
VideoMME (w sub.)
Video Understanding
VideoMME (with sub.) - comprehensive video understanding benchmark with subtitles.
VideoMME (w/o sub.)
Video Understanding
VideoMME (without sub.) - comprehensive video understanding benchmark without subtitles.
VideoMMMU
Video Understanding
VideoMMMU - multi-modal multi-discipline video understanding benchmark.
MVBench
Video Understanding
MVBench - comprehensive multi-modal video understanding benchmark.
Artificial Analysis Intelligence Index
Composite
Artificial Analysis Intelligence Index v4.1 - composite benchmark across reasoning, knowledge, ma...
ParseBench Mean
Document Understanding
ParseBench Mean - document parsing benchmark mean score across text content and formatting. Pipel...
ParseBench Text Content
Document Understanding
ParseBench Text Content - text content extraction accuracy from documents.
ParseBench Text Formatting
Document Understanding
ParseBench Text Formatting - text formatting fidelity in document parsing.
DynaMath
Vision & Language
DynaMath - dynamic mathematical reasoning in visual contexts.
VlmsAreBlind
Vision & Language
VlmsAreBlind - benchmark testing visual perception capabilities that reveal blind spots in VLMs.
MMStar
Vision & Language
MMStar - comprehensive multimodal benchmark evaluating fine-grained visual reasoning.
OCRBench
Document Understanding
OCRBench - comprehensive OCR capability evaluation benchmark.
ERQA
Spatial Intelligence
ERQA - embodied/spatial reasoning question answering benchmark.
V-Star
Vision & Language
V* - visual agent benchmark evaluating visual search and grounding capabilities.
AndroidWorld
General Agents
AndroidWorld - visual agent benchmark for Android device task completion.
AgentWorldBench MCP
General Agents
AgentWorldBench MCP domain - open-ended rubric evaluation (5 dimensions: Format, Factuality, Cons...
AgentWorldBench Search
General Agents
AgentWorldBench Search domain - open-ended rubric evaluation for language world models simulating...
AgentWorldBench Terminal
General Agents
AgentWorldBench Terminal domain - open-ended rubric evaluation for language world models simulati...
AgentWorldBench SWE
General Agents
AgentWorldBench SWE domain - open-ended rubric evaluation for language world models simulating so...
AgentWorldBench Android
General Agents
AgentWorldBench Android domain - open-ended rubric evaluation for language world models simulatin...
AgentWorldBench Web
General Agents
AgentWorldBench Web domain - open-ended rubric evaluation for language world models simulating we...
AgentWorldBench OS
General Agents
AgentWorldBench OS domain - open-ended rubric evaluation for language world models simulating ope...
AgentWorldBench Overall
General Agents
AgentWorldBench Overall - mean across all 7 agent interaction domains (MCP, Search, Terminal, SWE...
Terminal-Bench 2.1 (Terminus-2)
Agentic Coding
Terminal-Bench 2.1 evaluated with Harbor/Terminus-2 framework, parser=json, temp=1.0, top_p=1.0, ...
Terminal-Bench 2.1 (Claude Code)
Agentic Coding
Terminal-Bench 2.1 evaluated using Claude Code 2.1.126, parser=json, temp=1.0, top_p=1.0, max_new...
SWE Atlas - QnA
Agentic Coding
SWE Atlas QnA - agentic coding subtask; mini SWE agent harness, temp=1.0, top_p=0.95, 128K contex...
SWE Atlas - RF
Agentic Coding
SWE Atlas RF - agentic coding subtask; mini SWE agent harness, temp=1.0, top_p=0.95, 128K context...
SWE Atlas - TW
Agentic Coding
SWE Atlas TW - agentic coding subtask; mini SWE agent harness, temp=1.0, top_p=0.95, 128K context...
GDPVal
General Agents
GDPVal - agentic benchmark for evaluating general task completion.
TauBench V3 Airline
General Agents
TauBench V3 - Airline domain. Agentic tool use benchmark for airline customer service.
TauBench V3 Retail
General Agents
TauBench V3 - Retail domain. Agentic tool use benchmark for retail customer service.
TauBench V3 Telecom
General Agents
TauBench V3 - Telecom domain. Agentic tool use benchmark for telecom customer service.
TauBench V3 Banking
General Agents
TauBench V3 - Banking domain. Agentic tool use benchmark for banking customer service.
TauBench V3 Average
General Agents
TauBench V3 - Average across all domains (Airline, Retail, Telecom, Banking).
BrowseComp
General Agents
BrowseComp - web browsing and comprehension agentic benchmark.
Vals.ai Financial Agent 1.1 (without web search)
General Agents
Vals.ai Financial Agent 1.1 - without web search mode. Financial agentic benchmark.
Vals.ai Financial Agent 1.1 (with web search)
General Agents
Vals.ai Financial Agent 1.1 - with web search mode. Financial agentic benchmark.
IOI 2025
STEM & Reasoning
IOI 2025 - International Olympiad in Informatics 2025 competitive programming score.
IMOAnswerBench (with tools)
STEM & Reasoning
IMOAnswerBench - with tools variant. International Mathematical Olympiad answer evaluation benchm...
Apex-Shortlist (no tools)
STEM & Reasoning
Apex-Shortlist - no tools variant. Competitive math benchmark without tool use.
Apex-Shortlist (with tools)
STEM & Reasoning
Apex-Shortlist - with tools variant. Competitive math benchmark with tool use.
SciCode (subtask)
STEM & Reasoning
SciCode (subtask) - scientific code generation benchmark, subtask-level scoring.
HLE (with tools)
STEM & Reasoning
HLE (with tools) - Humanity's Last Exam with tool use variant.
CritPt (no tools)
STEM & Reasoning
CritPt - critical point identification benchmark, no tools variant.
OmniScience Accuracy
Knowledge
OmniScience Accuracy - scientific knowledge accuracy benchmark.
OmniScience Non-Hallucination
Knowledge
OmniScience Non-Hallucination - scientific knowledge non-hallucination rate benchmark.
IFBench (prompt loose)
Instruction Following
IFBench (prompt loose) - instruction following benchmark, loose prompt evaluation.
Multi-Challenge
Instruction Following
Multi-Challenge - multi-turn instruction following benchmark by ScaleAI.
Longbench v2 (≤ 1M)
Long Context
Longbench v2 (≤ 1M) - long context understanding benchmark, up to 1M tokens.
MMLU-ProX
Multilingual
MMLU-ProX - multilingual extended MMLU benchmark, average across en/de/fr/es/it/ja/zh/hi/pt/ko.
WMT24++ (en→xx)
Multilingual
WMT24++ (en→xx) - machine translation benchmark, English to other languages.
DeepSWE 1.1
Agentic Coding
DeepSWE 1.1 - agentic coding benchmark evaluated with Claude Code harness at temp=1.0, top_p=0.95...
QwenSWEBench
Agentic Coding
QwenSWEBench - in-house coding benchmark for evaluating software engineering capabilities. Evalua...
CoWorkBench
General Agents
CoWorkBench - in-house cowork benchmark for evaluating long-horizon tasks across computer science...
JobBench
General Agents
JobBench - in-house benchmark for evaluating professional job task completion.
Agents' Last Exam
General Agents
Agents' Last Exam - frontier agentic tasks benchmark. Reports Pass@1 and Score (aggregate).
OSWorld-Verified
General Agents
OSWorld-Verified - computer use benchmark evaluating autonomous GUI operation on verified tasks a...
WebArena-Verified
General Agents
WebArena-Verified - browser use benchmark evaluating autonomous web navigation and task completio...
RecreationBench
General Agents
RecreationBench - in-house long-horizon application-recreation benchmark evaluating hybrid-agent ...
SWE-MM
Agentic Coding
SWE-MM - multimodal software engineering benchmark. Evaluated on Claude Code harness using public...
Vision2Web
Vision & Language
Vision2Web - visual web development benchmark. Scores averaged across frontend, webpage, and webs...
MathVision
Vision & Language
MathVision - visual math problem solving benchmark. Reports Without CI and With CI (Consistency I...
BabyVision
Vision & Language
BabyVision - general visual reasoning benchmark. Reports Without CI and With CI (Consistency Infe...
Terminal Bench 2.1
Agentic Coding
Terminal Bench 2.1 - agentic coding and terminal use benchmark. Evaluated with Claude Code (avg@1...
NL2Repo-Bench
Agentic Coding
NL2Repo-Bench - repository generation from natural language. Evaluated with Claude Code harness. ...
FrontierSWE
Agentic Coding
FrontierSWE - agentic coding benchmark. MEAN@5 scores from the official FrontierSWE leaderboard. ...
MLS-Bench-Lite
Agentic Coding
MLS-Bench-Lite - agentic coding benchmark. Evaluated with Claude Code using a 5-hour timeout and ...
PaperBench
Agentic Coding
PaperBench - evaluated in the BasicAgent setting under Code-Dev mode, judged by Claude Opus 4.6, ...
AndroidBench
Agentic Coding
AndroidBench - evaluated on the 95-task public subset, reporting avg@3 scores.
QwenQoderBench
Agentic Coding
QwenQoderBench - in-house coding benchmark to evaluate user experience on Qoder. Evaluated with C...
QwenReactBench
Agentic Coding
QwenReactBench - in-house React project building benchmark using Claude Code harness. Bilingual (...
QwenSVGBench
Agentic Coding
QwenSVGBench - in-house SVG code generation benchmark; bilingual (EN/CN), auto-render + multimoda...
WorkSpaceBench
General Agents
WorkSpaceBench - in-house benchmark for evaluating workspace productivity tasks.
SkillsBench
General Agents
SkillsBench - evaluated on the public SkillsBench v1.1 benchmark across 87 tasks, reporting the a...
Automation-Bench
General Agents
Automation-Bench - evaluated on the 600-task public subset. Pass@1 score.
Toolathlon Verified
General Agents
Toolathlon Verified - agentic tool use benchmark. Pass@1 score.
IFBench
Instruction Following
IFBench - instruction following benchmark. General score from Qwen3.8 model card.
$OneMillion-Bench
General Capabilities
$OneMillion-Bench - expert score. Evaluated using gemini-3.1-pro-preview.
HealthBench
General Capabilities
HealthBench - healthcare-related benchmark evaluating medical knowledge and reasoning.
PLawBench
General Capabilities
PLawBench - professional law benchmark. Evaluated using gemini-3.1-pro-preview.
PRBench-Legal
General Capabilities
PRBench-Legal - professional reasoning benchmark for legal domain. Evaluated using gemini-3.1-pro...
PRBench-Finance
General Capabilities
PRBench-Finance - professional reasoning benchmark for finance domain. Evaluated using gemini-3.1...
MRCR v2 256K (8-needle)
Long Context
MRCR v2 256K (8-needle) - Multi-Ring Curve Retrieval benchmark v2, 256K context, 8-needle setting.
IFEval
Instruction Following
IFEval - Instruction Following Evaluation benchmark for measuring LLM instruction-following capab...
BFCL-V4
General Agents
BFCL-V4 - Berkeley Function Calling Leaderboard version 4, evaluating function calling and tool u...
TAU2-Bench
General Agents
TAU2-Bench - agentic tool use benchmark for customer service scenarios. Airline domain evaluated ...
MMMLU
Multilingual
MMMLU - Multilingual MMLU benchmark evaluating knowledge across multiple languages.
INCLUDE
Multilingual
INCLUDE - multilingual benchmark for evaluating language coverage across diverse regions.
Global PIQA
Multilingual
Global PIQA - physical reasoning benchmark across multiple languages and cultures.
MAXIFE
Multilingual
MAXIFE - multilingual instruction following benchmark. Reports accuracy on English + multilingual...
We-Math
Vision & Language
We-Math - visual mathematical reasoning benchmark evaluating step-by-step problem solving.
ZEROBench
Vision & Language
ZEROBench - visual reasoning puzzle benchmark. Very challenging, scores typically in single digits.
MMLongBench-Doc
Document Understanding
MMLongBench-Doc - long document understanding benchmark evaluating comprehension and retrieval ov...
LingoQA
Spatial Intelligence
LingoQA - spatial and linguistic reasoning benchmark for visual grounding.
Hypersim
Spatial Intelligence
Hypersim - 3D spatial understanding benchmark using synthetic indoor scenes.
Nuscene
Spatial Intelligence
Nuscene - 3D spatial understanding benchmark using autonomous driving scenes.
MMVU
Video Understanding
MMVU - multi-modal video understanding benchmark evaluating comprehension across diverse video ty...
ScreenSpot Pro
General Agents
ScreenSpot Pro - visual agent benchmark for GUI screen element grounding and interaction.
TIR-Bench
Vision & Language
TIR-Bench - tool-integrated reasoning benchmark for visual mathematical problem solving. Reports ...
SLAKE
Vision & Language
SLAKE - medical visual question answering benchmark using radiology images.
PMC-VQA
Vision & Language
PMC-VQA - medical visual question answering benchmark using PubMed Central biomedical images.
MedXpertQA-MM
Vision & Language
MedXpertQA-MM - multimodal medical expert question answering benchmark evaluating clinical reason...
Ifstruct V1
Instruction Following
Ifstruct V1 - instruction following structured evaluation benchmark by LiquidAI.
DeepSearch QA
General Agents
DeepSearch QA - full-task agentic benchmark measuring ability to work within scaffolds, write and...
WildClawBench
Agentic Coding
WildClawBench - agentic coding evaluation benchmark. Overall score from internlm/WildClawBench.
Gaia2
General Agents
Gaia2 - general AI assistant benchmark measuring multi-step agentic task completion.
Beam128K
Long Context
Beam128K - long context retrieval and understanding benchmark at 128K tokens.
MBCT
Safety
MBCT - Meta Biological CTF benchmark evaluating biological knowledge and wet-lab debugging capabi...
HPCT
Safety
HPCT - High-Performance CTF benchmark evaluating cyber preparedness capabilities.
VCT
Safety
VCT - Vulnerability CTF benchmark evaluating cyber security preparedness capabilities.
WMDP (Bio)
Safety
WMDP (Bio) - Weapons of Mass Destruction Prohibition (Biology) benchmark evaluating biological kn...
WMDP (Chem)
Safety
WMDP (Chem) - Weapons of Mass Destruction Prohibition (Chemistry) benchmark evaluating chemical k...
AIME 2025 (with tools)
STEM & Reasoning
AIME 2025 (with tools) - American Invitational Mathematics Examination 2025 with tool use variant.
Lab Bench (ProtocolQA)
Safety
Lab Bench (ProtocolQA) - wet-lab protocol question answering benchmark evaluating practical lab k...
GDPVal-AA v2
General Agents
GDPVal-AA v2 - Artificial Analysis GDPVal v2 agentic benchmark for evaluating general task comple...
ProgramBench
Agentic Coding
ProgramBench - agentic coding benchmark evaluated with Claude-Code 2.1.156, 200 instances, temp=1...
SWE-Marathon
Agentic Coding
SWE-Marathon - long-horizon software engineering benchmark. Evaluated by Abundant AI with 1M cont...
Terminal-Bench 2.1 (Best Reported Harness)
Agentic Coding
Terminal-Bench 2.1 (Best Reported Harness) - best score across all evaluation harnesses (Terminus...
BrowseComp-zh
General Agents
BrowseComp-zh - Chinese web browsing and comprehension agentic benchmark.
CodeForces
STEM & Reasoning
CodeForces - competitive programming rating evaluated on custom query set.
FullStackBench en
Agentic Coding
FullStackBench en - English full-stack code generation benchmark.
FullStackBench zh
Agentic Coding
FullStackBench zh - Chinese full-stack code generation benchmark.
SUNRGBD
Spatial Intelligence
SUNRGBD - 3D spatial understanding benchmark using RGB-D indoor scenes.
OCRBench
Document Understanding
OCRBench - comprehensive OCR capability evaluation benchmark.
HLE with search
STEM & Reasoning
HLE with search - Humanity's Last Exam with web search access variant.
BigBench Extra Hard
STEM & Reasoning
BigBench Extra Hard - extended and more challenging version of BigBench.
CoVoST
Multilingual
CoVoST - Conversational Voice Translation benchmark for speech-to-text translation across multipl...
FLEURS
Multilingual
FLEURS - Few-shot Learning Evaluation of Universal Representations of Speech. Lower is better (er...
MRCR v2 128K (8-needle)
Long Context
MRCR v2 128K (8-needle) - Multi-Ring Curve Retrieval benchmark v2, 128K context, 8-needle setting.
Cybergym
General Agents
Cybergym - agentic cybersecurity benchmark evaluating autonomous hacking and security task capabi...
DeepSWE
Agentic Coding
DeepSWE - agentic coding benchmark from datacurve/deep-swe dataset evaluating software engineerin...
DSBench-FullStack
Agentic Coding
DSBench-FullStack - internal full-stack development test set for evaluating coding agents on full...
DSBench-Hard
Agentic Coding
DSBench-Hard - internal test set of difficult coding-agent problems for evaluating agentic coding...
Aider
Agentic Coding
Aider - coding benchmark based on the Aider coding assistant tool
Frontier-Bench v0.1
Agentic Coding
Frontier-Bench v0.1 - frontier coding benchmark evaluating advanced agentic coding capabilities.
Kimi Code Bench V2
Agentic Coding
In-house benchmark by Moonshot AI designed to evaluate coding agents on realistic tasks. It has d...
Program Bench
Agentic Coding
Evaluates code-generation agents by asking them to recreate a program's behavior from only a comp...
Kimi Claw 24/7 Bench
Agentic
In-house benchmark by Moonshot AI for evaluating long-horizon agentic performance in persistent, ...
MCPMark-Verified
Agentic
A human-verified edition of MCPMark, a benchmark for evaluating MCP tool use across five real ser...
LHTB Solved
Agentic Coding
Long-Horizon-Terminal-Bench (LHTB) Solved metric from IntelligenceLab. Evaluates long-horizon ter...
BrowseComp Agent Swarm
General Agents
BrowseComp evaluated with agent swarm mode - multiple coordinated sub-agents performing web brows...
AIME 2024
STEM & Reasoning
AIME 2024 (American Invitational Mathematics Examination) - math competition problems evaluated p...
AIME 2025
STEM & Reasoning
AIME 2025 (American Invitational Mathematics Examination) - math competition problems evaluated p...
MMLU
Knowledge
MMLU (Massive Multitask Language Understanding) - a benchmark of multitask accuracy across 57 tasks.
Tau-Bench Retail
General Agents
τ-Bench Retail - function calling ability benchmark in a retail domain. Measures tool use with de...
Tau-Bench Airline
General Agents
τ-Bench Airline - function calling ability benchmark in an airline domain. Measures tool use with...
HealthBench Hard
General Capabilities
HealthBench Hard - a challenging subset of HealthBench conversations testing realistic health con...
HealthBench Consensus
General Capabilities
HealthBench Consensus - a subset of HealthBench validated by the consensus of multiple physicians.
Aider Polyglot
Agentic Coding
Aider Polyglot - coding benchmark evaluating model's ability to edit code across multiple program...
AIME 2024 (with tools)
STEM & Reasoning
AIME 2024 (with tools) - American Invitational Mathematics Examination 2024 with tool use variant.
GPQA Diamond (with tools)
STEM & Reasoning
GPQA Diamond (with tools) - Google-Proof Q&A benchmark for graduate-level scientific reasoning wi...
CodeForces (with tools)
STEM & Reasoning
CodeForces (with tools) - competitive programming rating evaluated with terminal tool access simi...
BoolQ
Reasoning
BoolQ is a QA task where each example comprises a short passage and a yes/no question. 0-shot.
PIQA
Reasoning
Physical Interaction QA (PIQA) - commonsense reasoning about physical world. 0-shot.
SocialIQA
Reasoning
Social IQA - commonsense reasoning about social situations. 0-shot.
Natural Questions
Knowledge
Google Natural Questions - open-domain QA using real user queries from Google Search. 5-shot.
ARC-c
Reasoning
AI2 Reasoning Challenge (Challenge set) - grade-school science questions. 25-shot.
ARC-e
Reasoning
AI2 Reasoning Challenge (Easy set) - grade-school science questions. 0-shot.
WinoGrande
Reasoning
WinoGrande - large-scale commonsense reasoning with Winograd schema-like problems. 5-shot.
BIG-Bench Hard
Reasoning
BIG-Bench Hard (BBH) - a subset of 23 challenging BIG-Bench tasks. few-shot.
AGIEval
Reasoning
AGIEval - human-level standardized tests (SAT, LSAT, GRE, etc.). 3-5-shot.
HumanEval
Code
HumanEval - Python code generation from function signatures and docstrings. 0-shot.
Global-MMLU-Lite
Multilingual
Global-MMLU-Lite - multilingual MMLU subset by Cohere For AI.
IndicGenBench
Multilingual
IndicGenBench - multilingual generation benchmark for Indic languages.
DocVQA
Multimodal
Document Visual Question Answering - understanding text in document images. (val split)
InfoVQA
Multimodal
Infographic VQA - answering questions about infographics. (val split)
TextVQA
Multimodal
TextVQA - visual question answering requiring reading text in images. (val split)
SpatialSense VQA
Multimodal
SpatialSense VQA - spatial reasoning in visual question answering.
OSWorld 2.0 (Binary)
General Agents
OSWorld 2.0 - Binary score. The binary score is the percentage of tasks that receive the full tas...
OSWorld 2.0 (Partial)
General Agents
OSWorld 2.0 - Partial score. The partial score aggregates the partial rewards obtained across all...
Agents' Last Exam (Score)
General Agents
Agents' Last Exam - Score (aggregate). Reports the aggregate score rather than Pass@1.
Terminal-Bench 3.0
Agentic Coding
Terminal-Bench 3.0 agentic terminal tasks, evaluated with the Claude Code 2.1.207 harness (reason...
ExploitGym (2h)
Cybersecurity
ExploitGym exploitation benchmark, single-run Pass@1 on 869 tasks under a 2-hour timeout budget (...
ExploitGym (6h)
Cybersecurity
ExploitGym exploitation benchmark, single-run Pass@1 on 869 tasks under a 6-hour timeout budget (...
ExploitBench
Cybersecurity
Exploitation benchmark: average coverage score over 41 tasks across 3 revisions (union of capabil...
Z.ai Code Bench
Agentic Coding
Z.AI's in-house coding benchmark for the GLM series. GLM-5.3 shows a 50% relative improvement ove...
Harbor-Index
Agentic Coding
Agentic coding index benchmark reported in the Tencent Hy4 preview model card benchmark appendix.
Hy-Backend 2.0 (Internal)
Agentic Coding
Tencent internal (Hy team) backend engineering benchmark, reported in the Hy4 preview model card ...
Hy-SWE Max Verified (Internal)
Agentic Coding
Tencent internal 300-task agentic coding benchmark; Claude Code scaffold, 200-turn budget, 16 CPU...
Hy-CompanyBench V2 (Internal)
General Agents
Tencent internal (Hy team) agentic search / company-work benchmark, reported in the Hy4 preview m...
Hy-LifeSearch (Internal)
General Agents
Tencent internal (Hy team) agentic search benchmark, reported in the Hy4 preview model card appen...
Hy-BrowseComp-Pro2 (Internal)
General Agents
Tencent internal (Hy team) hard web-browsing benchmark, reported in the Hy4 preview model card ap...
E-Bench (Internal)
General Agents
Tencent internal (Hy team) office/enterprise-work benchmark, reported in the Hy4 preview model ca...
E-Bench-Code (Internal)
Agentic Coding
Tencent internal (Hy team) enterprise coding benchmark, reported in the Hy4 preview model card ap...
Hy-FinAgentBench (Internal)
Finance
Tencent internal (Hy team) financial-agent benchmark, reported in the Hy4 preview model card appe...
Hy-FinmodelBench v2 (Internal)
Finance
Tencent internal (Hy team) financial-modelling benchmark, reported in the Hy4 preview model card ...
BioMysteryBench
STEM & Reasoning
Biology research-agent benchmark (99-task initial release); per model card footnote evaluated wit...
SUPERChem
STEM & Reasoning
Chemistry reasoning benchmark reported in the Tencent Hy4 preview model card appendix.
ArXivMath
STEM & Reasoning
Math reasoning benchmark over arXiv-style problems, reported in the Tencent Hy4 preview model car...
HorizonMath (pass@4)
STEM & Reasoning
HorizonMath - long-horizon math reasoning, pass@4 protocol, reported in the Tencent Hy4 preview m...
MathArena Apex 2025
STEM & Reasoning
MathArena Apex 2025 - hardest math contest problems of the 2025 MathArena set.
BrokenArXiv
STEM & Reasoning
Math reasoning benchmark reported in the Tencent Hy4 preview model card appendix.
Workspace Bench
Gaokao 2026
LCB-Pro 25Q2 (Easy)
STEM & Reasoning
LiveCodeBench Pro 25Q2 (Easy) - competitive programming evaluation, easy difficulty subset of LCB...
LCB-Pro 25Q2 (Medium)
STEM & Reasoning
LiveCodeBench Pro 25Q2 (Medium) - competitive programming evaluation, medium difficulty subset of...
NoLiMa
Long Context
NoLiMa - long context benchmark evaluating retrieval and reasoning without explicit long-context ...
LongBenchPro
Long Context
LongBenchPro - long context understanding benchmark with professional-level tasks.
Multi-IF
Instruction Following
Multi-IF - multi-turn instruction following benchmark evaluating complex multi-constraint instruc...
MATH-500
Math
MATH-500 - 500 competition-level mathematics problems, pass@1 without tools.
Terminal-Bench 4.0
Agentic Coding
Terminal-Bench 4.0 agentic terminal tasks. DeepSeek-V4.1-Flash evaluated with the Minimal mode of...
SEC-Bench Pro
General Agents
Security exploitation agent benchmark; DeepSeek-V4.1-Flash evaluated with the Claude Code harness...
LiveBench 241125
Reasoning
LiveBench (2024-11-25 version) - contamination-free LLM benchmark with regularly updated tasks ac...
CFEval
Agentic Coding
CFEval - Qwen's CodeForces-style competitive programming evaluation; score is an Elo-like rating ...
Arena-Hard v2
Instruction Following
Arena-Hard v2 - instruction-following chat benchmark; win rate evaluated by GPT-4.1 for reproduci...
WritingBench
Instruction Following
WritingBench - benchmark evaluating fine-grained writing capabilities with verifiable constraints.
BFCL-V3
General Agents
BFCL-V3 - Berkeley Function Calling Leaderboard version 3, evaluating function calling and tool use.
TauBench V2 Retail
General Agents
tau2-bench - Retail domain. Agentic tool use benchmark for retail customer service scenarios.
TauBench V2 Airline
General Agents
tau2-bench - Airline domain. Agentic tool use benchmark for airline customer service scenarios.
TauBench V2 Telecom
General Agents
tau2-bench - Telecom domain. Agentic tool use benchmark for telecom customer service scenarios.
Creative Writing v3
Instruction Following
Creative Writing v3 - Qwen's creative writing quality evaluation benchmark.