Opus-4.8

Anthropic

Parameters

—

Architecture

—

Released

—

License

—

Claude

About

Opus-4.8 - proprietary model by Anthropic, used as comparison model in DeepSeek-V4-Pro-0813 benchmark table.

Benchmark Scores

Benchmark Score Date
Terminal Bench 2.1
coding_agent
93.79%
31.07.2026
NL2Repo
coding_agent
91.38%
31.07.2026
Cybergym
general_agent
89.88%
31.07.2026
DeepSWE
coding_agent
79.78%
31.07.2026
Toolathlon Verified
general_agent
96.61%
31.07.2026
Agents' Last Exam
general_agent
71.23%
31.07.2026
Automation-Bench
general_agent
37.27%
31.07.2026
DSBench-FullStack
coding_agent
86.07%
31.07.2026
DSBench-Hard
coding_agent
100.00%
31.07.2026
Humanity's Last Exam
stem_reasoning
90.34%
13.08.2026
HLE (with tools)
stem_reasoning
86.02%
13.08.2026
Terminal-Bench 3.0
coding_agent
55.00%
28.08.2026
DeepSWE 1.1
coding_agent
75.53%
28.08.2026
ProgramBench
coding_agent
18.14%
28.08.2026
FrontierSWE
coding_agent
63.34%
28.08.2026
SWE-Marathon
coding_agent
97.55%
28.08.2026
PostTrainBench
coding_agent
67.58%
28.08.2026
ExploitGym (2h)
cybersecurity
32.67%
28.08.2026
ExploitGym (6h)
cybersecurity
35.21%
28.08.2026
ExploitBench
cybersecurity
29.10%
28.08.2026
GDPVal-AA v2
general_agent
86.73%
28.08.2026
ApexBench
multimodal_agent
100.00%
08.09.2026
Chartography
multimodal_agent
4.79%
08.09.2026
ZEROBench
vision_language
67.39%
08.09.2026

Related Models