Claude Opus 4.6 vs Claude Opus 4.7
New ComparisonScores are fetched live from the current database on every visit — nothing is cached.
Benchmark Scores Comparison
Claude Opus 4.6
Claude Opus 4.7
Claude Opus 4.6
Score: 15.0
Claude Opus 4.7
Score: 15.0
Coverage matrix
2 models × 58 benchmarks — hover a benchmark name for its category
Shading = score relative to the best score among the selected models:
Low
Mid
High
Missing (—)
The totals row shows how many of the 58 benchmarks each model has scores for.
| Benchmark | Claude Opus 4.6 Anthropic | Claude Opus 4.7 Anthropic |
|---|---|---|
| SVG-Bench | — | 62.3 |
| GPQA Diamond | 91.3 | — |
| LiveCodeBench v6 | 88.8 | — |
| SWE-bench Pro | 53.4 | 64.3 |
| MathVision | 71.2 | — |
| AIME 26 | 96.7 | — |
| LiveSQLBench | — | 41 |
| AgentWorldBench Terminal | 57.51 | — |
| Toolathlon Verified | 47.2 | — |
| OJBench | 60.3 | — |
| AgentWorldBench Overall | 57.8 | — |
| GDPVal | — | 79.8 |
| OfficeQA Pro | — | 43.6 |
| DeepSearch QA | 91.3 | — |
| SWE Atlas - TW | — | 38.2 |
| CL-bench | — | 22.9 |
| LOCA-Bench (256k) | — | 57 |
| KernelBench Hard | — | 30.7 |
| AgentWorldBench Web | 51.42 | — |
| SWE Atlas - QnA | — | 45.2 |
| IMO 2025 | — | 17.4 |
| MMMU-Pro | 73.9 | 77 |
| PostTrainBench | — | 42.4 |
| BabyVision | 14.8 | — |
| Claw-Eval Avg | 82.4 | 71.6 |
| VideoMMMU | — | 83 |
| Apex-Agents | 33 | 37.2 |
| BankerToolBench | — | 81.3 |
| OSWorld-Verified | 72.7 | 82.8 |
| AgentWorldBench Search | 29.3 | — |
| HLE (with tools) | 53 | — |
| CharXiv (RQ) | 69.1 | — |
| BrowseComp | 83.7 | 79.3 |
| NL2Repo-Bench | — | 56.3 |
| DRACO | — | 77.7 |
| MCPMark | 56.7 | — |
| SWE-bench Multilingual | 77.8 | — |
| PaperBench | — | 58.5 |
| YC-Bench | — | 2200000 |
| SWE-bench Verified | 80.8 | 87.6 |
| IMOAnswerBench | 75.3 | — |
| Humanity's Last Exam | 40 | — |
| Terminal-Bench 2.0 | 65.4 | — |
| SpreadSheetBench-v1 | — | 88.5 |
| VIBE-V2 | — | 55.9 |
| OmniDocBench 1.5 | — | 89.3 |
| HMMT Feb 26 | 96.2 | — |
| AgentWorldBench OS | 70.2 | — |
| AgentWorldBench MCP | 69.9 | — |
| Terminal Bench 2.1 | — | 66.1 |
| AgentWorldBench SWE | 64.55 | — |
| SciCode (subtask) | 51.9 | — |
| USAMO 2026 | — | 52.8 |
| Claw-Eval Pass^3 | 70.4 | — |
| MCP-Atlas | — | 77 |
| V-Star | 86.4 | — |
| AgentWorldBench Android | 61.74 | — |
| SWE-efficiency | — | 42.2 |
| Covered | 34/58 | 31/58 |
Fully comparable benchmarks
Scored by all 2 selected models (7)
| Benchmark | Category | Claude Opus 4.6 | Claude Opus 4.7 |
|---|---|---|---|
| SWE-bench Pro | coding_agent | 53.4 | 64.3 |
| MMMU-Pro | vision_language | 73.9 | 77 |
| Claw-Eval Avg | coding_agent | 82.4 | 71.6 |
| Apex-Agents | general_agent | 33 | 37.2 |
| OSWorld-Verified | general_agent | 72.7 | 82.8 |
| BrowseComp | general_agent | 83.7 | 79.3 |
| SWE-bench Verified | coding_agent | 80.8 | 87.6 |
Partial coverage benchmarks
Scored by at least one but not all selected models (51) — missing scores show as —
| Benchmark | Category | Claude Opus 4.6 | Claude Opus 4.7 |
|---|---|---|---|
| SVG-Bench 1/2 models | coding_agent | — | 62.3 |
| GPQA Diamond 1/2 models | stem_reasoning | 91.3 | — |
| LiveCodeBench v6 1/2 models | stem_reasoning | 88.8 | — |
| MathVision 1/2 models | vision_language | 71.2 | — |
| AIME 26 1/2 models | stem_reasoning | 96.7 | — |
| LiveSQLBench 1/2 models | coding_agent | — | 41 |
| AgentWorldBench Terminal 1/2 models | general_agent | 57.51 | — |
| Toolathlon Verified 1/2 models | general_agent | 47.2 | — |
| OJBench 1/2 models | stem_reasoning | 60.3 | — |
| AgentWorldBench Overall 1/2 models | general_agent | 57.8 | — |
| GDPVal 1/2 models | general_agent | — | 79.8 |
| OfficeQA Pro 1/2 models | general_agent | — | 43.6 |
| DeepSearch QA 1/2 models | general_agent | 91.3 | — |
| SWE Atlas - TW 1/2 models | coding_agent | — | 38.2 |
| CL-bench 1/2 models | coding_agent | — | 22.9 |
| LOCA-Bench (256k) 1/2 models | long_context | — | 57 |
| KernelBench Hard 1/2 models | coding_agent | — | 30.7 |
| AgentWorldBench Web 1/2 models | general_agent | 51.42 | — |
| SWE Atlas - QnA 1/2 models | coding_agent | — | 45.2 |
| IMO 2025 1/2 models | stem_reasoning | — | 17.4 |
| PostTrainBench 1/2 models | coding_agent | — | 42.4 |
| BabyVision 1/2 models | vision_language | 14.8 | — |
| VideoMMMU 1/2 models | video_understanding | — | 83 |
| BankerToolBench 1/2 models | general_agent | — | 81.3 |
| AgentWorldBench Search 1/2 models | general_agent | 29.3 | — |
| HLE (with tools) 1/2 models | stem_reasoning | 53 | — |
| CharXiv (RQ) 1/2 models | document_understanding | 69.1 | — |
| NL2Repo-Bench 1/2 models | coding_agent | — | 56.3 |
| DRACO 1/2 models | general_agent | — | 77.7 |
| MCPMark 1/2 models | general_agent | 56.7 | — |
| SWE-bench Multilingual 1/2 models | coding_agent | 77.8 | — |
| PaperBench 1/2 models | coding_agent | — | 58.5 |
| YC-Bench 1/2 models | general_agent | — | 2200000 |
| IMOAnswerBench 1/2 models | stem_reasoning | 75.3 | — |
| Humanity's Last Exam 1/2 models | stem_reasoning | 40 | — |
| Terminal-Bench 2.0 1/2 models | coding_agent | 65.4 | — |
| SpreadSheetBench-v1 1/2 models | general_agent | — | 88.5 |
| VIBE-V2 1/2 models | coding_agent | — | 55.9 |
| OmniDocBench 1.5 1/2 models | document_understanding | — | 89.3 |
| HMMT Feb 26 1/2 models | stem_reasoning | 96.2 | — |
| AgentWorldBench OS 1/2 models | general_agent | 70.2 | — |
| AgentWorldBench MCP 1/2 models | general_agent | 69.9 | — |
| Terminal Bench 2.1 1/2 models | coding_agent | — | 66.1 |
| AgentWorldBench SWE 1/2 models | general_agent | 64.55 | — |
| SciCode (subtask) 1/2 models | stem_reasoning | 51.9 | — |
| USAMO 2026 1/2 models | stem_reasoning | — | 52.8 |
| Claw-Eval Pass^3 1/2 models | coding_agent | 70.4 | — |
| MCP-Atlas 1/2 models | general_agent | — | 77 |
| V-Star 1/2 models | vision_language | 86.4 | — |
| AgentWorldBench Android 1/2 models | general_agent | 61.74 | — |
| SWE-efficiency 1/2 models | coding_agent | — | 42.2 |