Opus-4.8
Anthropic
Parameters
—
Architecture
—
Released
—
License
—
About
Opus-4.8 - proprietary model by Anthropic, used as comparison model in DeepSeek-V4-Pro-0813 benchmark table.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
Terminal Bench 2.1
coding_agent
|
93.79%
|
31.07.2026 |
|
NL2Repo
coding_agent
|
91.38%
|
31.07.2026 |
|
Cybergym
general_agent
|
89.88%
|
31.07.2026 |
|
DeepSWE
coding_agent
|
79.78%
|
31.07.2026 |
|
Toolathlon Verified
general_agent
|
96.61%
|
31.07.2026 |
|
Agents' Last Exam
general_agent
|
71.23%
|
31.07.2026 |
|
Automation-Bench
general_agent
|
37.27%
|
31.07.2026 |
|
DSBench-FullStack
coding_agent
|
86.07%
|
31.07.2026 |
|
DSBench-Hard
coding_agent
|
100.00%
|
31.07.2026 |
|
Humanity's Last Exam
stem_reasoning
|
90.34%
|
13.08.2026 |
|
HLE (with tools)
stem_reasoning
|
86.02%
|
13.08.2026 |
|
Terminal-Bench 3.0
coding_agent
|
55.00%
|
28.08.2026 |
|
DeepSWE 1.1
coding_agent
|
75.53%
|
28.08.2026 |
|
ProgramBench
coding_agent
|
18.14%
|
28.08.2026 |
|
FrontierSWE
coding_agent
|
63.34%
|
28.08.2026 |
|
SWE-Marathon
coding_agent
|
97.55%
|
28.08.2026 |
|
PostTrainBench
coding_agent
|
67.58%
|
28.08.2026 |
|
ExploitGym (2h)
cybersecurity
|
32.67%
|
28.08.2026 |
|
ExploitGym (6h)
cybersecurity
|
35.21%
|
28.08.2026 |
|
ExploitBench
cybersecurity
|
29.10%
|
28.08.2026 |
|
GDPVal-AA v2
general_agent
|
86.73%
|
28.08.2026 |
|
ApexBench
multimodal_agent
|
100.00%
|
08.09.2026 |
|
Chartography
multimodal_agent
|
4.79%
|
08.09.2026 |
|
ZEROBench
vision_language
|
67.39%
|
08.09.2026 |