LLM Benchmarks — Scores, Methodology & Live Rankings

What each benchmark measures, how models score, and the current rankings

Reset

SWE-bench Verified

Agentic Coding

SWE-bench Verified evaluates LLM-based agents on real-world GitHub issue resolution from popular ...

Active View →

SWE-bench Multilingual

Agentic Coding

SWE-bench Multilingual extends SWE-bench to multiple programming languages beyond Python.

Active View →

SWE-bench Pro

Agentic Coding

SWE-bench Pro - refined benchmark with corrected problematic tasks in the public set. Internal ag...

Active View →

Terminal-Bench 2.0

Agentic Coding

Terminal-Bench 2.0 - agentic coding and terminal use benchmark. Harbor/Terminus-2 harness; 3h tim...

Active View →

Claw-Eval Avg

Agentic Coding

Claw-Eval average score - agentic coding evaluation.

Active View →

Claw-Eval Pass^3

Agentic Coding

Claw-Eval Pass^3 - third-pass pass rate on the Claw-Eval agentic coding benchmark.

Active View →

SkillsBench Avg5

Agentic Coding

SkillsBench Avg5 - evaluated via OpenCode on 78 self-contained tasks (excluding API-dependent tas...

Active View →

QwenClawBench

Agentic Coding

QwenClawBench - internal real-user-distribution Claw agent benchmark (open-sourcing soon); temp=0...

Active View →

NL2Repo

Agentic Coding

NL2Repo - repository generation from natural language. Others evaluated via Claude Code (temp=1.0...

Active View →

QwenWebBench

Agentic Coding

QwenWebBench - internal front-end code generation benchmark; bilingual (EN/CN), 7 categories (Web...

Active View →

TAU3-Bench

General Agents

TAU3-Bench - agentic tool use benchmark. Uses official user model (gpt-5.2, low reasoning effort)...

Active View →

VITA-Bench

General Agents

VITA-Bench - agentic benchmark. Avg subdomain scores; using claude-4-sonnet as judger.

Active View →

DeepPlanning

General Agents

DeepPlanning - agentic planning benchmark.

Active View →

Tool Decathlon

General Agents

Tool Decathlon - agentic tool use benchmark.

Active View →

MCPMark

General Agents

MCPMark - GitHub MCP v0.30.3; Playwright responses truncated at 32K tokens.

Active View →

MCP-Atlas

General Agents

MCP-Atlas - public set score; gemini-2.5-pro judger.

Active View →

WideSearch

General Agents

WideSearch - agentic web search benchmark.

Active View →

MMLU-Pro

Knowledge

MMLU-Pro - extended MMLU covering 14 categories with complex reasoning questions.

Active View →

MMLU-Redux

Knowledge

MMLU-Redux - cleaned and verified subset of MMLU.

Active View →

SuperGPQA

Knowledge

SuperGPQA - graduate-level multi-domain QA benchmark.

Active View →

C-Eval

Knowledge

C-Eval - comprehensive Chinese evaluation benchmark across 52 subjects.

Active View →

GPQA Diamond

STEM & Reasoning

GPQA Diamond - Google-Proof Q&A benchmark for graduate-level scientific reasoning.

Active View →

Humanity's Last Exam

STEM & Reasoning

HLE - Humanity's Last Exam - extremely difficult expert-level questions across domains.

Active View →

LiveCodeBench v6

STEM & Reasoning

LiveCodeBench v6 - competitive programming evaluation with latest problems.

Active View →

HMMT Feb 25

STEM & Reasoning

HMMT Feb 25 - Harvard-MIT Mathematics Tournament February 2025.

Active View →

HMMT Nov 25

STEM & Reasoning

HMMT Nov 25 - Harvard-MIT Mathematics Tournament November 2025.

Active View →

HMMT Feb 26

STEM & Reasoning

HMMT Feb 26 - Harvard-MIT Mathematics Tournament February 2026.

Active View →

IMOAnswerBench

STEM & Reasoning

IMOAnswerBench - International Mathematical Olympiad answer evaluation benchmark.

Active View →

AIME 26

STEM & Reasoning

AIME 26 - American Invitational Mathematics Examination 2026 (I & II).

Active View →

MMMU

Vision & Language

MMMU - massive multi-discipline multimodal understanding and reasoning benchmark for college-leve...

Active View →

MMMU-Pro

Vision & Language

MMMU-Pro - enhanced MMMU with harder questions and refined evaluation.

Active View →

MathVista (mini)

Vision & Language

MathVista (mini) - mathematical reasoning in visual contexts.

Active View →

ZEROBench_sub

Vision & Language

ZEROBench subset - visual reasoning puzzle benchmark.

Active View →

RealWorldQA

Vision & Language

RealWorldQA - visual question answering on real-world images.

Active View →

MMBench EN-DEV v1.1

Vision & Language

MMBench-EN-DEV v1.1 - multimodal benchmark English dev set.

Active View →

SimpleVQA

Vision & Language

SimpleVQA - simple visual question answering benchmark.

Active View →

HallusionBench

Vision & Language

HallusionBench - benchmark for language and visual hallucination.

Active View →

OmniDocBench 1.5

Document Understanding

OmniDocBench 1.5 - document understanding and parsing benchmark.

Active View →

CharXiv (RQ)

Document Understanding

CharXiv (RQ) - reasoning questions over arXiv figures and tables.

Active View →

CC-OCR

Document Understanding

CC-OCR - comprehensive OCR benchmark across diverse content.

Active View →

AI2D_TEST

Document Understanding

AI2D_TEST - grade-school science diagram understanding benchmark.

Active View →

RefCOCO (avg)

Spatial Intelligence

RefCOCO (avg) - referring expression comprehension benchmark, average score.

Active View →

ODInW13

Spatial Intelligence

ODInW13 - object detection in the wild across 13 domains.

Active View →

EmbSpatialBench

Spatial Intelligence

EmbSpatialBench - embodied spatial intelligence benchmark.

Active View →

RefSpatialBench

Spatial Intelligence

RefSpatialBench - referring spatial reasoning benchmark.

Active View →

VideoMME (w sub.)

Video Understanding

VideoMME (with sub.) - comprehensive video understanding benchmark with subtitles.

Active View →

VideoMME (w/o sub.)

Video Understanding

VideoMME (without sub.) - comprehensive video understanding benchmark without subtitles.

Active View →

VideoMMMU

Video Understanding

VideoMMMU - multi-modal multi-discipline video understanding benchmark.

Active View →

MLVU

Video Understanding

MLVU - multi-task long video understanding benchmark.

Active View →

MVBench

Video Understanding

MVBench - comprehensive multi-modal video understanding benchmark.

Active View →

LVBench

Video Understanding

LVBench - long video understanding benchmark.

Active View →

Artificial Analysis Intelligence Index

Composite

Artificial Analysis Intelligence Index v4.1 - composite benchmark across reasoning, knowledge, ma...

Active View →

ParseBench Mean

Document Understanding

ParseBench Mean - document parsing benchmark mean score across text content and formatting. Pipel...

Active View →

ParseBench Text Content

Document Understanding

ParseBench Text Content - text content extraction accuracy from documents.

Active View →

ParseBench Text Formatting

Document Understanding

ParseBench Text Formatting - text formatting fidelity in document parsing.

Active View →

DynaMath

Vision & Language

DynaMath - dynamic mathematical reasoning in visual contexts.

Active View →

VlmsAreBlind

Vision & Language

VlmsAreBlind - benchmark testing visual perception capabilities that reveal blind spots in VLMs.

Active View →

MMStar

Vision & Language

MMStar - comprehensive multimodal benchmark evaluating fine-grained visual reasoning.

Active View →

OCRBench

Document Understanding

OCRBench - comprehensive OCR capability evaluation benchmark.

Active View →

ERQA

Spatial Intelligence

ERQA - embodied/spatial reasoning question answering benchmark.

Active View →

CountBench

Spatial Intelligence

CountBench - visual object counting benchmark.

Active View →

V-Star

Vision & Language

V* - visual agent benchmark evaluating visual search and grounding capabilities.

Active View →

AndroidWorld

General Agents

AndroidWorld - visual agent benchmark for Android device task completion.

Active View →

AgentWorldBench MCP

General Agents

AgentWorldBench MCP domain - open-ended rubric evaluation (5 dimensions: Format, Factuality, Cons...

Active View →

AgentWorldBench Search

General Agents

AgentWorldBench Search domain - open-ended rubric evaluation for language world models simulating...

Active View →

AgentWorldBench Terminal

General Agents

AgentWorldBench Terminal domain - open-ended rubric evaluation for language world models simulati...

Active View →

AgentWorldBench SWE

General Agents

AgentWorldBench SWE domain - open-ended rubric evaluation for language world models simulating so...

Active View →

AgentWorldBench Android

General Agents

AgentWorldBench Android domain - open-ended rubric evaluation for language world models simulatin...

Active View →

AgentWorldBench Web

General Agents

AgentWorldBench Web domain - open-ended rubric evaluation for language world models simulating we...

Active View →

AgentWorldBench OS

General Agents

AgentWorldBench OS domain - open-ended rubric evaluation for language world models simulating ope...

Active View →

AgentWorldBench Overall

General Agents

AgentWorldBench Overall - mean across all 7 agent interaction domains (MCP, Search, Terminal, SWE...

Active View →

Terminal-Bench 2.1 (Terminus-2)

Agentic Coding

Terminal-Bench 2.1 evaluated with Harbor/Terminus-2 framework, parser=json, temp=1.0, top_p=1.0, ...

Active View →

Terminal-Bench 2.1 (Claude Code)

Agentic Coding

Terminal-Bench 2.1 evaluated using Claude Code 2.1.126, parser=json, temp=1.0, top_p=1.0, max_new...

Active View →

SWE Atlas - QnA

Agentic Coding

SWE Atlas QnA - agentic coding subtask; mini SWE agent harness, temp=1.0, top_p=0.95, 128K contex...

Active View →

SWE Atlas - RF

Agentic Coding

SWE Atlas RF - agentic coding subtask; mini SWE agent harness, temp=1.0, top_p=0.95, 128K context...

Active View →

SWE Atlas - TW

Agentic Coding

SWE Atlas TW - agentic coding subtask; mini SWE agent harness, temp=1.0, top_p=0.95, 128K context...

Active View →

GDPVal

General Agents

GDPVal - agentic benchmark for evaluating general task completion.

Active View →

ProfBench (Search)

General Agents

ProfBench (Search) - agentic search benchmark.

Active View →

PinchBench

General Agents

PinchBench - agentic benchmark.

Active View →

TauBench V3 Airline

General Agents

TauBench V3 - Airline domain. Agentic tool use benchmark for airline customer service.

Active View →

TauBench V3 Retail

General Agents

TauBench V3 - Retail domain. Agentic tool use benchmark for retail customer service.

Active View →

TauBench V3 Telecom

General Agents

TauBench V3 - Telecom domain. Agentic tool use benchmark for telecom customer service.

Active View →

TauBench V3 Banking

General Agents

TauBench V3 - Banking domain. Agentic tool use benchmark for banking customer service.

Active View →

TauBench V3 Average

General Agents

TauBench V3 - Average across all domains (Airline, Retail, Telecom, Banking).

Active View →

BrowseComp

General Agents

BrowseComp - web browsing and comprehension agentic benchmark.

Active View →

Vals.ai Financial Agent 1.1 (without web search)

General Agents

Vals.ai Financial Agent 1.1 - without web search mode. Financial agentic benchmark.

Active View →

Vals.ai Financial Agent 1.1 (with web search)

General Agents

Vals.ai Financial Agent 1.1 - with web search mode. Financial agentic benchmark.

Active View →

IOI 2025

STEM & Reasoning

IOI 2025 - International Olympiad in Informatics 2025 competitive programming score.

Active View →

IMOAnswerBench (with tools)

STEM & Reasoning

IMOAnswerBench - with tools variant. International Mathematical Olympiad answer evaluation benchm...

Active View →

Apex-Shortlist (no tools)

STEM & Reasoning

Apex-Shortlist - no tools variant. Competitive math benchmark without tool use.

Active View →

Apex-Shortlist (with tools)

STEM & Reasoning

Apex-Shortlist - with tools variant. Competitive math benchmark with tool use.

Active View →

SciCode (subtask)

STEM & Reasoning

SciCode (subtask) - scientific code generation benchmark, subtask-level scoring.

Active View →

HLE (with tools)

STEM & Reasoning

HLE (with tools) - Humanity's Last Exam with tool use variant.

Active View →

CritPt (no tools)

STEM & Reasoning

CritPt - critical point identification benchmark, no tools variant.

Active View →

OmniScience Accuracy

Knowledge

OmniScience Accuracy - scientific knowledge accuracy benchmark.

Active View →

OmniScience Non-Hallucination

Knowledge

OmniScience Non-Hallucination - scientific knowledge non-hallucination rate benchmark.

Active View →

IFBench (prompt loose)

Instruction Following

IFBench (prompt loose) - instruction following benchmark, loose prompt evaluation.

Active View →

Multi-Challenge

Instruction Following

Multi-Challenge - multi-turn instruction following benchmark by ScaleAI.

Active View →

AA-LCR

Long Context

AA-LCR - Artificial Analysis Long Context Retrieval benchmark.

Active View →

RULER (1M)

Long Context

RULER (1M) - long context retrieval benchmark at 1M tokens.

Active View →

Longbench v2 (≤ 1M)

Long Context

Longbench v2 (≤ 1M) - long context understanding benchmark, up to 1M tokens.

Active View →

MMLU-ProX

Multilingual

MMLU-ProX - multilingual extended MMLU benchmark, average across en/de/fr/es/it/ja/zh/hi/pt/ko.

Active View →

WMT24++ (en→xx)

Multilingual

WMT24++ (en→xx) - machine translation benchmark, English to other languages.

Active View →

DeepSWE 1.1

Agentic Coding

DeepSWE 1.1 - agentic coding benchmark evaluated with Claude Code harness at temp=1.0, top_p=0.95...

Active View →

QwenSWEBench

Agentic Coding

QwenSWEBench - in-house coding benchmark for evaluating software engineering capabilities. Evalua...

Active View →

CoWorkBench

General Agents

CoWorkBench - in-house cowork benchmark for evaluating long-horizon tasks across computer science...

Active View →

JobBench

General Agents

JobBench - in-house benchmark for evaluating professional job task completion.

Active View →

Agents' Last Exam

General Agents

Agents' Last Exam - frontier agentic tasks benchmark. Reports Pass@1 and Score (aggregate).

Active View →

OSWorld-Verified

General Agents

OSWorld-Verified - computer use benchmark evaluating autonomous GUI operation on verified tasks a...

Active View →

WebArena-Verified

General Agents

WebArena-Verified - browser use benchmark evaluating autonomous web navigation and task completio...

Active View →

RecreationBench

General Agents

RecreationBench - in-house long-horizon application-recreation benchmark evaluating hybrid-agent ...

Active View →

SWE-MM

Agentic Coding

SWE-MM - multimodal software engineering benchmark. Evaluated on Claude Code harness using public...

Active View →

Vision2Web

Vision & Language

Vision2Web - visual web development benchmark. Scores averaged across frontend, webpage, and webs...

Active View →

MathVision

Vision & Language

MathVision - visual math problem solving benchmark. Reports Without CI and With CI (Consistency I...

Active View →

BabyVision

Vision & Language

BabyVision - general visual reasoning benchmark. Reports Without CI and With CI (Consistency Infe...

Active View →

Terminal Bench 2.1

Agentic Coding

Terminal Bench 2.1 - agentic coding and terminal use benchmark. Evaluated with Claude Code (avg@1...

Active View →

NL2Repo-Bench

Agentic Coding

NL2Repo-Bench - repository generation from natural language. Evaluated with Claude Code harness. ...

Active View →

FrontierSWE

Agentic Coding

FrontierSWE - agentic coding benchmark. MEAN@5 scores from the official FrontierSWE leaderboard. ...

Active View →

MLS-Bench-Lite

Agentic Coding

MLS-Bench-Lite - agentic coding benchmark. Evaluated with Claude Code using a 5-hour timeout and ...

Active View →

PaperBench

Agentic Coding

PaperBench - evaluated in the BasicAgent setting under Code-Dev mode, judged by Claude Opus 4.6, ...

Active View →

AndroidBench

Agentic Coding

AndroidBench - evaluated on the 95-task public subset, reporting avg@3 scores.

Active View →

QwenQoderBench

Agentic Coding

QwenQoderBench - in-house coding benchmark to evaluate user experience on Qoder. Evaluated with C...

Active View →

QwenReactBench

Agentic Coding

QwenReactBench - in-house React project building benchmark using Claude Code harness. Bilingual (...

Active View →

QwenSVGBench

Agentic Coding

QwenSVGBench - in-house SVG code generation benchmark; bilingual (EN/CN), auto-render + multimoda...

Active View →

WorkSpaceBench

General Agents

WorkSpaceBench - in-house benchmark for evaluating workspace productivity tasks.

Active View →

SkillsBench

General Agents

SkillsBench - evaluated on the public SkillsBench v1.1 benchmark across 87 tasks, reporting the a...

Active View →

Automation-Bench

General Agents

Automation-Bench - evaluated on the 600-task public subset. Pass@1 score.

Active View →

Toolathlon Verified

General Agents

Toolathlon Verified - agentic tool use benchmark. Pass@1 score.

Active View →

IFBench

Instruction Following

IFBench - instruction following benchmark. General score from Qwen3.8 model card.

Active View →

$OneMillion-Bench

General Capabilities

$OneMillion-Bench - expert score. Evaluated using gemini-3.1-pro-preview.

Active View →

HealthBench

General Capabilities

HealthBench - healthcare-related benchmark evaluating medical knowledge and reasoning.

Active View →

PLawBench

General Capabilities

PLawBench - professional law benchmark. Evaluated using gemini-3.1-pro-preview.

Active View →

PRBench-Legal

General Capabilities

PRBench-Legal - professional reasoning benchmark for legal domain. Evaluated using gemini-3.1-pro...

Active View →

PRBench-Finance

General Capabilities

PRBench-Finance - professional reasoning benchmark for finance domain. Evaluated using gemini-3.1...

Active View →

MRCR v2 256K (8-needle)

Long Context

MRCR v2 256K (8-needle) - Multi-Ring Curve Retrieval benchmark v2, 256K context, 8-needle setting.

Active View →

LongBench v2

Long Context

LongBench v2 - long context understanding benchmark.

Active View →

IFEval

Instruction Following

IFEval - Instruction Following Evaluation benchmark for measuring LLM instruction-following capab...

Active View →

OJBench

STEM & Reasoning

OJBench - competitive programming and reasoning benchmark.

Active View →

BFCL-V4

General Agents

BFCL-V4 - Berkeley Function Calling Leaderboard version 4, evaluating function calling and tool u...

Active View →

TAU2-Bench

General Agents

TAU2-Bench - agentic tool use benchmark for customer service scenarios. Airline domain evaluated ...

Active View →

MMMLU

Multilingual

MMMLU - Multilingual MMLU benchmark evaluating knowledge across multiple languages.

Active View →

NOVA-63

Multilingual

NOVA-63 - multilingual benchmark covering 63 languages.

Active View →

INCLUDE

Multilingual

INCLUDE - multilingual benchmark for evaluating language coverage across diverse regions.

Active View →

Global PIQA

Multilingual

Global PIQA - physical reasoning benchmark across multiple languages and cultures.

Active View →

PolyMATH

Multilingual

PolyMATH - multilingual mathematical reasoning benchmark.

Active View →

MAXIFE

Multilingual

MAXIFE - multilingual instruction following benchmark. Reports accuracy on English + multilingual...

Active View →

We-Math

Vision & Language

We-Math - visual mathematical reasoning benchmark evaluating step-by-step problem solving.

Active View →

ZEROBench

Vision & Language

ZEROBench - visual reasoning puzzle benchmark. Very challenging, scores typically in single digits.

Active View →

MMLongBench-Doc

Document Understanding

MMLongBench-Doc - long document understanding benchmark evaluating comprehension and retrieval ov...

Active View →

LingoQA

Spatial Intelligence

LingoQA - spatial and linguistic reasoning benchmark for visual grounding.

Active View →

Hypersim

Spatial Intelligence

Hypersim - 3D spatial understanding benchmark using synthetic indoor scenes.

Active View →

Nuscene

Spatial Intelligence

Nuscene - 3D spatial understanding benchmark using autonomous driving scenes.

Active View →

MMVU

Video Understanding

MMVU - multi-modal video understanding benchmark evaluating comprehension across diverse video ty...

Active View →

ScreenSpot Pro

General Agents

ScreenSpot Pro - visual agent benchmark for GUI screen element grounding and interaction.

Active View →

TIR-Bench

Vision & Language

TIR-Bench - tool-integrated reasoning benchmark for visual mathematical problem solving. Reports ...

Active View →

SLAKE

Vision & Language

SLAKE - medical visual question answering benchmark using radiology images.

Active View →

PMC-VQA

Vision & Language

PMC-VQA - medical visual question answering benchmark using PubMed Central biomedical images.

Active View →

MedXpertQA-MM

Vision & Language

MedXpertQA-MM - multimodal medical expert question answering benchmark evaluating clinical reason...

Active View →

Ifstruct V1

Instruction Following

Ifstruct V1 - instruction following structured evaluation benchmark by LiquidAI.

Active View →

DeepSearch QA

General Agents

DeepSearch QA - full-task agentic benchmark measuring ability to work within scaffolds, write and...

Active View →

WildClawBench

Agentic Coding

WildClawBench - agentic coding evaluation benchmark. Overall score from internlm/WildClawBench.

Active View →

Gaia2

General Agents

Gaia2 - general AI assistant benchmark measuring multi-step agentic task completion.

Active View →

Beam128K

Long Context

Beam128K - long context retrieval and understanding benchmark at 128K tokens.

Active View →

MBCT

Safety

MBCT - Meta Biological CTF benchmark evaluating biological knowledge and wet-lab debugging capabi...

Active View →

HPCT

Safety

HPCT - High-Performance CTF benchmark evaluating cyber preparedness capabilities.

Active View →

VCT

Safety

VCT - Vulnerability CTF benchmark evaluating cyber security preparedness capabilities.

Active View →

WMDP (Bio)

Safety

WMDP (Bio) - Weapons of Mass Destruction Prohibition (Biology) benchmark evaluating biological kn...

Active View →

WMDP (Chem)

Safety

WMDP (Chem) - Weapons of Mass Destruction Prohibition (Chemistry) benchmark evaluating chemical k...

Active View →

AIME 2025 (with tools)

STEM & Reasoning

AIME 2025 (with tools) - American Invitational Mathematics Examination 2025 with tool use variant.

Active View →

Lab Bench (ProtocolQA)

Safety

Lab Bench (ProtocolQA) - wet-lab protocol question answering benchmark evaluating practical lab k...

Active View →

GDPVal-AA v2

General Agents

GDPVal-AA v2 - Artificial Analysis GDPVal v2 agentic benchmark for evaluating general task comple...

Active View →

Apex-Agents

General Agents

Active View →

SpreadSheetBench-v1

General Agents

Active View →

YC-Bench

General Agents

Active View →

SWE-efficiency

Agentic Coding

Active View →

LiveSQLBench

Agentic Coding

Active View →

CL-bench

Agentic Coding

Active View →

VIBE-V2

Agentic Coding

Active View →

SVG-Bench

Agentic Coding

Active View →

PostTrainBench

Agentic Coding

Active View →

KernelBench Hard

Agentic Coding

Active View →

DRACO

General Agents

Active View →

BankerToolBench

General Agents

Active View →

OfficeQA Pro

General Agents

Active View →

IMO 2025

STEM & Reasoning

Active View →

USAMO 2026

STEM & Reasoning

Active View →

ProgramBench

Agentic Coding

ProgramBench - agentic coding benchmark evaluated with Claude-Code 2.1.156, 200 instances, temp=1...

Active View →

SWE-Marathon

Agentic Coding

SWE-Marathon - long-horizon software engineering benchmark. Evaluated by Abundant AI with 1M cont...

Active View →

Terminal-Bench 2.1 (Best Reported Harness)

Agentic Coding

Terminal-Bench 2.1 (Best Reported Harness) - best score across all evaluation harnesses (Terminus...

Active View →

BrowseComp-zh

General Agents

BrowseComp-zh - Chinese web browsing and comprehension agentic benchmark.

Active View →

Seal-0

General Agents

Seal-0 - search agent agentic benchmark.

Active View →

CodeForces

STEM & Reasoning

CodeForces - competitive programming rating evaluated on custom query set.

Active View →

FullStackBench en

Agentic Coding

FullStackBench en - English full-stack code generation benchmark.

Active View →

FullStackBench zh

Agentic Coding

FullStackBench zh - Chinese full-stack code generation benchmark.

Active View →

SUNRGBD

Spatial Intelligence

SUNRGBD - 3D spatial understanding benchmark using RGB-D indoor scenes.

Active View →

OCRBench

Document Understanding

OCRBench - comprehensive OCR capability evaluation benchmark.

Active View →

HLE with search

STEM & Reasoning

HLE with search - Humanity's Last Exam with web search access variant.

Active View →

BigBench Extra Hard

STEM & Reasoning

BigBench Extra Hard - extended and more challenging version of BigBench.

Active View →

CoVoST

Multilingual

CoVoST - Conversational Voice Translation benchmark for speech-to-text translation across multipl...

Active View →

FLEURS

Multilingual

FLEURS - Few-shot Learning Evaluation of Universal Representations of Speech. Lower is better (er...

Active View →

MRCR v2 128K (8-needle)

Long Context

MRCR v2 128K (8-needle) - Multi-Ring Curve Retrieval benchmark v2, 128K context, 8-needle setting.

Active View →

Cybergym

General Agents

Cybergym - agentic cybersecurity benchmark evaluating autonomous hacking and security task capabi...

Active View →

DeepSWE

Agentic Coding

DeepSWE - agentic coding benchmark from datacurve/deep-swe dataset evaluating software engineerin...

Active View →

DSBench-FullStack

Agentic Coding

DSBench-FullStack - internal full-stack development test set for evaluating coding agents on full...

Active View →

DSBench-Hard

Agentic Coding

DSBench-Hard - internal test set of difficult coding-agent problems for evaluating agentic coding...

Active View →

Aider

Agentic Coding

Aider - coding benchmark based on the Aider coding assistant tool

Active View →

Frontier-Bench v0.1

Agentic Coding

Frontier-Bench v0.1 - frontier coding benchmark evaluating advanced agentic coding capabilities.

Active View →

Kimi Code Bench V2

Agentic Coding

In-house benchmark by Moonshot AI designed to evaluate coding agents on realistic tasks. It has d...

Active View →

Program Bench

Agentic Coding

Evaluates code-generation agents by asking them to recreate a program's behavior from only a comp...

Active View →

Kimi Claw 24/7 Bench

Agentic

In-house benchmark by Moonshot AI for evaluating long-horizon agentic performance in persistent, ...

Active View →

MCPMark-Verified

Agentic

A human-verified edition of MCPMark, a benchmark for evaluating MCP tool use across five real ser...

Active View →

LHTB Solved

Agentic Coding

Long-Horizon-Terminal-Bench (LHTB) Solved metric from IntelligenceLab. Evaluates long-horizon ter...

Active View →

BrowseComp Agent Swarm

General Agents

BrowseComp evaluated with agent swarm mode - multiple coordinated sub-agents performing web brows...

Active View →

AIME 2024

STEM & Reasoning

AIME 2024 (American Invitational Mathematics Examination) - math competition problems evaluated p...

Active View →

AIME 2025

STEM & Reasoning

AIME 2025 (American Invitational Mathematics Examination) - math competition problems evaluated p...

Active View →

MMLU

Knowledge

MMLU (Massive Multitask Language Understanding) - a benchmark of multitask accuracy across 57 tasks.

Active View →

Tau-Bench Retail

General Agents

τ-Bench Retail - function calling ability benchmark in a retail domain. Measures tool use with de...

Active View →

Tau-Bench Airline

General Agents

τ-Bench Airline - function calling ability benchmark in an airline domain. Measures tool use with...

Active View →

HealthBench Hard

General Capabilities

HealthBench Hard - a challenging subset of HealthBench conversations testing realistic health con...

Active View →

HealthBench Consensus

General Capabilities

HealthBench Consensus - a subset of HealthBench validated by the consensus of multiple physicians.

Active View →

Aider Polyglot

Agentic Coding

Aider Polyglot - coding benchmark evaluating model's ability to edit code across multiple program...

Active View →

AIME 2024 (with tools)

STEM & Reasoning

AIME 2024 (with tools) - American Invitational Mathematics Examination 2024 with tool use variant.

Active View →

GPQA Diamond (with tools)

STEM & Reasoning

GPQA Diamond (with tools) - Google-Proof Q&A benchmark for graduate-level scientific reasoning wi...

Active View →

CodeForces (with tools)

STEM & Reasoning

CodeForces (with tools) - competitive programming rating evaluated with terminal tool access simi...

Active View →

HellaSwag

Reasoning

HellaSwag is a benchmark for commonsense NLI. 10-shot.

Active View →

BoolQ

Reasoning

BoolQ is a QA task where each example comprises a short passage and a yes/no question. 0-shot.

Active View →

PIQA

Reasoning

Physical Interaction QA (PIQA) - commonsense reasoning about physical world. 0-shot.

Active View →

SocialIQA

Reasoning

Social IQA - commonsense reasoning about social situations. 0-shot.

Active View →

TriviaQA

Knowledge

TriviaQA - reading comprehension with trivia questions. 5-shot.

Active View →

Natural Questions

Knowledge

Google Natural Questions - open-domain QA using real user queries from Google Search. 5-shot.

Active View →

ARC-c

Reasoning

AI2 Reasoning Challenge (Challenge set) - grade-school science questions. 25-shot.

Active View →

ARC-e

Reasoning

AI2 Reasoning Challenge (Easy set) - grade-school science questions. 0-shot.

Active View →

WinoGrande

Reasoning

WinoGrande - large-scale commonsense reasoning with Winograd schema-like problems. 5-shot.

Active View →

BIG-Bench Hard

Reasoning

BIG-Bench Hard (BBH) - a subset of 23 challenging BIG-Bench tasks. few-shot.

Active View →

DROP

Reasoning

DROP - discrete reasoning over paragraphs. 1-shot.

Active View →

MMLU Pro COT

Knowledge

MMLU-Pro with chain-of-thought reasoning. 5-shot.

Active View →

AGIEval

Reasoning

AGIEval - human-level standardized tests (SAT, LSAT, GRE, etc.). 3-5-shot.

Active View →

MATH

Math

MATH benchmark - competition-level mathematics problems. 4-shot.

Active View →

GSM8K

Math

Grade School Math 8K - grade-school math word problems. 8-shot.

Active View →

GPQA

Reasoning

Google-Proof Q&A - graduate-level science questions. 5-shot.

Active View →

MBPP

Code

Mostly Basic Python Problems - Python coding tasks. 3-shot.

Active View →

HumanEval

Code

HumanEval - Python code generation from function signatures and docstrings. 0-shot.

Active View →

MGSM

Multilingual

Multilingual Grade School Math - GSM8K translated to 10 languages.

Active View →

Global-MMLU-Lite

Multilingual

Global-MMLU-Lite - multilingual MMLU subset by Cohere For AI.

Active View →

FloRes

Multilingual

FloRes - machine translation benchmark covering 200+ languages.

Active View →

XQuAD

Multilingual

Cross-lingual Question Answering Dataset - QA in 10+ languages.

Active View →

ECLeKTic

Multilingual

ECLeKTic - multilingual knowledge translation benchmark.

Active View →

IndicGenBench

Multilingual

IndicGenBench - multilingual generation benchmark for Indic languages.

Active View →

COCOcap

Multimodal

COCO Captioning - image captioning on MS-COCO dataset.

Active View →

DocVQA

Multimodal

Document Visual Question Answering - understanding text in document images. (val split)

Active View →

InfoVQA

Multimodal

Infographic VQA - answering questions about infographics. (val split)

Active View →

TextVQA

Multimodal

TextVQA - visual question answering requiring reading text in images. (val split)

Active View →

ReMI

Multimodal

ReMI - referring to images in multimodal context.

Active View →

ChartQA

Multimodal

ChartQA - answering questions about charts and plots.

Active View →

VQAv2

Multimodal

Visual Question Answering v2 - general image QA.

Active View →

BLINK

Multimodal

BLINK - multimodal benchmark for visual perception tasks.

Active View →

OKVQA

Multimodal

OK-VQA - outside knowledge visual question answering.

Active View →

TallyQA

Multimodal

TallyQA - visual counting question answering.

Active View →

SpatialSense VQA

Multimodal

SpatialSense VQA - spatial reasoning in visual question answering.

Active View →

CountBenchQA

Multimodal

CountBenchQA - visual counting question answering benchmark.

Active View →

GPQA Chain-of-Thought

Reasoning

GPQA with chain-of-thought reasoning.

Active View →

OSWorld 2.0 (Binary)

General Agents

OSWorld 2.0 - Binary score. The binary score is the percentage of tasks that receive the full tas...

Active View →

OSWorld 2.0 (Partial)

General Agents

OSWorld 2.0 - Partial score. The partial score aggregates the partial rewards obtained across all...

Active View →

Agents' Last Exam (Score)

General Agents

Agents' Last Exam - Score (aggregate). Reports the aggregate score rather than Pass@1.

Active View →

Terminal-Bench 3.0

Agentic Coding

Terminal-Bench 3.0 agentic terminal tasks, evaluated with the Claude Code 2.1.207 harness (reason...

Active View →

ExploitGym (2h)

Cybersecurity

ExploitGym exploitation benchmark, single-run Pass@1 on 869 tasks under a 2-hour timeout budget (...

Active View →

ExploitGym (6h)

Cybersecurity

ExploitGym exploitation benchmark, single-run Pass@1 on 869 tasks under a 6-hour timeout budget (...

Active View →

ExploitBench

Cybersecurity

Exploitation benchmark: average coverage score over 41 tasks across 3 revisions (union of capabil...

Active View →

Z.ai Code Bench

Agentic Coding

Z.AI's in-house coding benchmark for the GLM series. GLM-5.3 shows a 50% relative improvement ove...

Active View →

Harbor-Index

Agentic Coding

Agentic coding index benchmark reported in the Tencent Hy4 preview model card benchmark appendix.

Active View →

Hy-Backend 2.0 (Internal)

Agentic Coding

Tencent internal (Hy team) backend engineering benchmark, reported in the Hy4 preview model card ...

Active View →

Hy-SWE Max Verified (Internal)

Agentic Coding

Tencent internal 300-task agentic coding benchmark; Claude Code scaffold, 200-turn budget, 16 CPU...

Active View →

Hy-CompanyBench V2 (Internal)

General Agents

Tencent internal (Hy team) agentic search / company-work benchmark, reported in the Hy4 preview m...

Active View →

Hy-LifeSearch (Internal)

General Agents

Tencent internal (Hy team) agentic search benchmark, reported in the Hy4 preview model card appen...

Active View →

Hy-BrowseComp-Pro2 (Internal)

General Agents

Tencent internal (Hy team) hard web-browsing benchmark, reported in the Hy4 preview model card ap...

Active View →

E-Bench (Internal)

General Agents

Tencent internal (Hy team) office/enterprise-work benchmark, reported in the Hy4 preview model ca...

Active View →

E-Bench-Code (Internal)

Agentic Coding

Tencent internal (Hy team) enterprise coding benchmark, reported in the Hy4 preview model card ap...

Active View →

Hy-FinAgentBench (Internal)

Finance

Tencent internal (Hy team) financial-agent benchmark, reported in the Hy4 preview model card appe...

Active View →

Hy-FinmodelBench v2 (Internal)

Finance

Tencent internal (Hy team) financial-modelling benchmark, reported in the Hy4 preview model card ...

Active View →

BioMysteryBench

STEM & Reasoning

Biology research-agent benchmark (99-task initial release); per model card footnote evaluated wit...

Active View →

SUPERChem

STEM & Reasoning

Chemistry reasoning benchmark reported in the Tencent Hy4 preview model card appendix.

Active View →

ArXivMath

STEM & Reasoning

Math reasoning benchmark over arXiv-style problems, reported in the Tencent Hy4 preview model car...

Active View →

HorizonMath (pass@4)

STEM & Reasoning

HorizonMath - long-horizon math reasoning, pass@4 protocol, reported in the Tencent Hy4 preview m...

Active View →

MathArena Apex 2025

STEM & Reasoning

MathArena Apex 2025 - hardest math contest problems of the 2025 MathArena set.

Active View →

BrokenArXiv

STEM & Reasoning

Math reasoning benchmark reported in the Tencent Hy4 preview model card appendix.

Active View →

ApexBench

Multimodal Agents

Active View →

Chartography

Multimodal Agents

Active View →

LCB-Pro 25Q2 (Easy)

STEM & Reasoning

LiveCodeBench Pro 25Q2 (Easy) - competitive programming evaluation, easy difficulty subset of LCB...

Active View →

LCB-Pro 25Q2 (Medium)

STEM & Reasoning

LiveCodeBench Pro 25Q2 (Medium) - competitive programming evaluation, medium difficulty subset of...

Active View →

NoLiMa

Long Context

NoLiMa - long context benchmark evaluating retrieval and reasoning without explicit long-context ...

Active View →

LongBenchPro

Long Context

LongBenchPro - long context understanding benchmark with professional-level tasks.

Active View →

Multi-IF

Instruction Following

Multi-IF - multi-turn instruction following benchmark evaluating complex multi-constraint instruc...

Active View →

MATH-500

Math

MATH-500 - 500 competition-level mathematics problems, pass@1 without tools.

Active View →

Terminal-Bench 4.0

Agentic Coding

Terminal-Bench 4.0 agentic terminal tasks. DeepSeek-V4.1-Flash evaluated with the Minimal mode of...

Active View →

SEC-Bench Pro

General Agents

Security exploitation agent benchmark; DeepSeek-V4.1-Flash evaluated with the Claude Code harness...

Active View →

LiveBench 241125

Reasoning

LiveBench (2024-11-25 version) - contamination-free LLM benchmark with regularly updated tasks ac...

Active View →

CFEval

Agentic Coding

CFEval - Qwen's CodeForces-style competitive programming evaluation; score is an Elo-like rating ...

Active View →

Arena-Hard v2

Instruction Following

Arena-Hard v2 - instruction-following chat benchmark; win rate evaluated by GPT-4.1 for reproduci...

Active View →

WritingBench

Instruction Following

WritingBench - benchmark evaluating fine-grained writing capabilities with verifiable constraints.

Active View →

BFCL-V3

General Agents

BFCL-V3 - Berkeley Function Calling Leaderboard version 3, evaluating function calling and tool use.

Active View →

TauBench V2 Retail

General Agents

tau2-bench - Retail domain. Agentic tool use benchmark for retail customer service scenarios.

Active View →

TauBench V2 Airline

General Agents

tau2-bench - Airline domain. Agentic tool use benchmark for airline customer service scenarios.

Active View →

TauBench V2 Telecom

General Agents

tau2-bench - Telecom domain. Agentic tool use benchmark for telecom customer service scenarios.

Active View →

Creative Writing v3

Instruction Following

Creative Writing v3 - Qwen's creative writing quality evaluation benchmark.

Active View →