Claude Opus 5

Anthropic

Parameters

—

Architecture

—

Released

—

License

—

Commercial Use Claude
Openness Index Score 25.0/100

About

Proprietary Anthropic frontier model; included as comparison model in the Tencent Hy4 preview benchmark appendix.

Benchmark Scores

Benchmark Score Date
SWE-bench Pro
coding_agent
99.00%
28.08.2026
DeepSWE
coding_agent
94.64%
28.08.2026
SWE Atlas - QnA
coding_agent
88.14%
28.08.2026
SWE Atlas - TW
coding_agent
100.00%
28.08.2026
SWE Atlas - RF
coding_agent
100.00%
28.08.2026
SWE-Marathon
coding_agent
100.00%
28.08.2026
Terminal-Bench 2.1 (Terminus-2)
coding_agent
96.90%
28.08.2026
NL2Repo
coding_agent
100.00%
28.08.2026
ProgramBench
coding_agent
52.98%
28.08.2026
PostTrainBench
coding_agent
74.74%
28.08.2026
Harbor-Index
coding_agent
100.00%
28.08.2026
Hy-Backend 2.0 (Internal)
coding_agent
60.26%
28.08.2026
Hy-SWE Max Verified (Internal)
coding_agent
100.00%
28.08.2026
Hy-CompanyBench V2 (Internal)
general_agent
100.00%
28.08.2026
WideSearch
general_agent
95.50%
28.08.2026
$OneMillion-Bench
general_capabilities
100.00%
28.08.2026
Hy-LifeSearch (Internal)
general_agent
70.20%
28.08.2026
Hy-BrowseComp-Pro2 (Internal)
general_agent
100.00%
28.08.2026
OfficeQA Pro
general_agent
100.00%
28.08.2026
MCP-Atlas
general_agent
100.00%
28.08.2026
Toolathlon Verified
general_agent
97.21%
28.08.2026
Apex-Agents
general_agent
100.00%
28.08.2026
SkillsBench Avg5
coding_agent
95.18%
28.08.2026
JobBench
general_agent
100.00%
28.08.2026
WorkSpaceBench
general_agent
100.00%
28.08.2026
GDPVal-AA v2
general_agent
100.00%
28.08.2026
Automation-Bench
general_agent
86.14%
28.08.2026
BankerToolBench
general_agent
100.00%
28.08.2026
E-Bench (Internal)
general_agent
91.28%
28.08.2026
E-Bench-Code (Internal)
coding_agent
96.84%
28.08.2026
Hy-FinAgentBench (Internal)
domain_finance
92.59%
28.08.2026
Hy-FinmodelBench v2 (Internal)
domain_finance
100.00%
28.08.2026
BioMysteryBench
stem_reasoning
94.51%
28.08.2026
HLE (with tools)
stem_reasoning
92.37%
28.08.2026
CritPt (no tools)
stem_reasoning
89.91%
28.08.2026
GPQA Diamond
stem_reasoning
98.87%
28.08.2026
Humanity's Last Exam
stem_reasoning
100.00%
28.08.2026
SUPERChem
stem_reasoning
100.00%
28.08.2026
ArXivMath
stem_reasoning
71.22%
28.08.2026
HorizonMath (pass@4)
stem_reasoning
25.28%
28.08.2026
MathArena Apex 2025
stem_reasoning
100.00%
28.08.2026
BrokenArXiv
stem_reasoning
100.00%
28.08.2026
DRACO
general_agent
100.00%
28.08.2026
SWE-bench Multilingual
coding_agent
100.00%
28.08.2026
Agents' Last Exam
general_agent
68.40%
28.08.2026

Related Models