Claude Opus 4.8
Anthropic
Parameters
—
Architecture
—
Released
—
License
—
About
No detailed description available for this model.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
AgentWorldBench MCP
general_agent
|
54.93
|
24.06.2026 |
|
AgentWorldBench Search
general_agent
|
83.12%
|
24.06.2026 |
|
AgentWorldBench Terminal
general_agent
|
100.00%
|
24.06.2026 |
|
AgentWorldBench SWE
general_agent
|
85.86%
|
24.06.2026 |
|
AgentWorldBench Android
general_agent
|
97.43%
|
24.06.2026 |
|
AgentWorldBench Web
general_agent
|
100.00%
|
24.06.2026 |
|
AgentWorldBench OS
general_agent
|
74.30%
|
24.06.2026 |
|
AgentWorldBench Overall
general_agent
|
83.16%
|
24.06.2026 |
|
Terminal Bench 2.1
coding_agent
|
93.35%
|
18.08.2026 |
|
SWE-bench Pro
coding_agent
|
86.50%
|
18.08.2026 |
|
DeepSWE 1.1
coding_agent
|
77.04%
|
18.08.2026 |
|
NL2Repo-Bench
coding_agent
|
100.00%
|
18.08.2026 |
|
FrontierSWE
coding_agent
|
69.26%
|
18.08.2026 |
|
MLS-Bench-Lite
coding_agent
|
69.40%
|
18.08.2026 |
|
PaperBench
coding_agent
|
79.65%
|
18.08.2026 |
|
AndroidBench
coding_agent
|
47.50%
|
18.08.2026 |
|
QwenSWEBench
coding_agent
|
93.78%
|
18.08.2026 |
|
QwenQoderBench
coding_agent
|
98.48%
|
18.08.2026 |
|
QwenReactBench
coding_agent
|
67.24%
|
18.08.2026 |
|
QwenSVGBench
coding_agent
|
57.53%
|
18.08.2026 |
|
CoWorkBench
general_agent
|
88.31%
|
18.08.2026 |
|
WorkSpaceBench
general_agent
|
51.19%
|
18.08.2026 |
|
JobBench
general_agent
|
57.58%
|
18.08.2026 |
|
SkillsBench
general_agent
|
79.56%
|
18.08.2026 |
|
Agents' Last Exam
general_agent
|
77.36%
|
18.08.2026 |
|
Automation-Bench
general_agent
|
37.27%
|
18.08.2026 |
|
Toolathlon Verified
general_agent
|
96.61%
|
18.08.2026 |
|
WideSearch
general_agent
|
73.78%
|
18.08.2026 |
|
HLE (with tools)
stem_reasoning
|
86.02%
|
18.08.2026 |
|
GPQA Diamond
stem_reasoning
|
95.65%
|
18.08.2026 |
|
Humanity's Last Exam
stem_reasoning
|
82.58%
|
18.08.2026 |
|
IFBench
instruction_following
|
65.72%
|
18.08.2026 |
|
$OneMillion-Bench
general_capabilities
|
41.80
|
18.08.2026 |
|
HealthBench
general_capabilities
|
55.93%
|
18.08.2026 |
|
PLawBench
general_capabilities
|
74.83%
|
18.08.2026 |
|
PRBench-Legal
general_capabilities
|
46.15%
|
18.08.2026 |
|
PRBench-Finance
general_capabilities
|
44.35%
|
18.08.2026 |
|
MRCR v2 256K (8-needle)
long_context
|
83.20
|
18.08.2026 |
|
LongBench v2
long_context
|
100.00%
|
18.08.2026 |
|
CritPt (no tools)
stem_reasoning
|
64.04%
|
17.06.2026 |
|
AIME 26
stem_reasoning
|
95.54%
|
17.06.2026 |
|
HMMT Nov 25
stem_reasoning
|
100.00%
|
17.06.2026 |
|
HMMT Feb 26
stem_reasoning
|
98.70%
|
17.06.2026 |
|
IMOAnswerBench
stem_reasoning
|
86.51%
|
17.06.2026 |
|
NL2Repo
coding_agent
|
91.38%
|
17.06.2026 |
|
ProgramBench
coding_agent
|
100.00%
|
17.06.2026 |
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
94.40%
|
17.06.2026 |
|
Terminal-Bench 2.1 (Best Reported Harness)
coding_agent
|
68.75%
|
17.06.2026 |
|
PostTrainBench
coding_agent
|
82.25%
|
17.06.2026 |
|
SWE-Marathon
coding_agent
|
51.02%
|
17.06.2026 |
|
MCP-Atlas
general_agent
|
89.19%
|
17.06.2026 |
|
Tool Decathlon
general_agent
|
100.00%
|
17.06.2026 |
|
Kimi Code Bench V2
coding_agent
|
91.16%
|
23.08.2026 |
|
Program Bench
coding_agent
|
74.52%
|
23.08.2026 |
|
Kimi Claw 24/7 Bench
agentic
|
75.76%
|
23.08.2026 |
|
MCPMark-Verified
agentic
|
17.91%
|
23.08.2026 |
|
GDPVal-AA v2
general_agent
|
86.40%
|
27.08.2026 |