GPT-5.6 Sol
OpenAI
Parameters
—
Architecture
—
Released
—
License
—
About
No detailed description available for this model.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
Terminal Bench 2.1
coding_agent
|
98.00%
|
18.08.2026 |
|
SWE-bench Pro
coding_agent
|
80.75%
|
18.08.2026 |
|
DeepSWE 1.1
coding_agent
|
98.19%
|
18.08.2026 |
|
MLS-Bench-Lite
coding_agent
|
84.05%
|
18.08.2026 |
|
PaperBench
coding_agent
|
95.99%
|
18.08.2026 |
|
AndroidBench
coding_agent
|
62.50%
|
18.08.2026 |
|
QwenSWEBench
coding_agent
|
65.41%
|
18.08.2026 |
|
QwenQoderBench
coding_agent
|
64.64%
|
18.08.2026 |
|
QwenReactBench
coding_agent
|
11.21%
|
18.08.2026 |
|
QwenSVGBench
coding_agent
|
100.00%
|
18.08.2026 |
|
CoWorkBench
general_agent
|
85.71%
|
18.08.2026 |
|
JobBench
general_agent
|
51.08%
|
18.08.2026 |
|
SkillsBench
general_agent
|
100.00%
|
18.08.2026 |
|
Agents' Last Exam
general_agent
|
94.34%
|
18.08.2026 |
|
Automation-Bench
general_agent
|
42.95%
|
18.08.2026 |
|
Toolathlon Verified
general_agent
|
94.01%
|
18.08.2026 |
|
HLE (with tools)
stem_reasoning
|
86.23%
|
18.08.2026 |
|
GPQA Diamond
stem_reasoning
|
99.62%
|
18.08.2026 |
|
Humanity's Last Exam
stem_reasoning
|
85.42%
|
18.08.2026 |
|
IFBench
instruction_following
|
83.19%
|
18.08.2026 |
|
$OneMillion-Bench
general_capabilities
|
45.63%
|
18.08.2026 |
|
HealthBench
general_capabilities
|
72.32%
|
18.08.2026 |
|
PLawBench
general_capabilities
|
93.71%
|
18.08.2026 |
|
PRBench-Legal
general_capabilities
|
100.00%
|
18.08.2026 |
|
PRBench-Finance
general_capabilities
|
75.65%
|
18.08.2026 |
|
MRCR v2 256K (8-needle)
long_context
|
100.00%
|
18.08.2026 |
|
LongBench v2
long_context
|
95.48%
|
18.08.2026 |
|
Terminal-Bench 3.0
coding_agent
|
100.00%
|
28.08.2026 |
|
SWE-Marathon
coding_agent
|
84.69%
|
28.08.2026 |
|
PostTrainBench
coding_agent
|
78.84%
|
28.08.2026 |
|
Cybergym
general_agent
|
90.89%
|
28.08.2026 |
|
ExploitGym (2h)
cybersecurity
|
100.00%
|
28.08.2026 |
|
ExploitGym (6h)
cybersecurity
|
100.00%
|
28.08.2026 |
|
ExploitBench
cybersecurity
|
97.20%
|
28.08.2026 |
|
GDPVal-AA v2
general_agent
|
94.48%
|
28.08.2026 |
|
SWE-bench Multilingual
coding_agent
|
81.78%
|
28.08.2026 |
|
DeepSWE
coding_agent
|
100.00%
|
28.08.2026 |
|
SWE Atlas - QnA
coding_agent
|
89.23%
|
28.08.2026 |
|
SWE Atlas - TW
coding_agent
|
70.30%
|
28.08.2026 |
|
SWE Atlas - RF
coding_agent
|
86.36%
|
28.08.2026 |
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
100.00%
|
28.08.2026 |
|
NL2Repo
coding_agent
|
71.54%
|
28.08.2026 |
|
Harbor-Index
coding_agent
|
74.33%
|
28.08.2026 |
|
Hy-Backend 2.0 (Internal)
coding_agent
|
100.00%
|
28.08.2026 |
|
Hy-SWE Max Verified (Internal)
coding_agent
|
98.58%
|
28.08.2026 |
|
Hy-CompanyBench V2 (Internal)
general_agent
|
95.10%
|
28.08.2026 |
|
WideSearch
general_agent
|
100.00%
|
28.08.2026 |
|
DRACO
general_agent
|
53.42%
|
28.08.2026 |
|
Hy-LifeSearch (Internal)
general_agent
|
100.00%
|
28.08.2026 |
|
Hy-BrowseComp-Pro2 (Internal)
general_agent
|
66.89%
|
28.08.2026 |
|
OfficeQA Pro
general_agent
|
96.93%
|
28.08.2026 |
|
MCP-Atlas
general_agent
|
95.62%
|
28.08.2026 |
|
Apex-Agents
general_agent
|
94.75%
|
28.08.2026 |
|
SkillsBench Avg5
coding_agent
|
93.26%
|
28.08.2026 |
|
BankerToolBench
general_agent
|
83.89%
|
28.08.2026 |
|
E-Bench (Internal)
general_agent
|
100.00%
|
28.08.2026 |
|
E-Bench-Code (Internal)
coding_agent
|
100.00%
|
28.08.2026 |
|
Hy-FinAgentBench (Internal)
domain_finance
|
100.00%
|
28.08.2026 |
|
Hy-FinmodelBench v2 (Internal)
domain_finance
|
98.13%
|
28.08.2026 |
|
BioMysteryBench
stem_reasoning
|
100.00%
|
28.08.2026 |
|
CritPt (no tools)
stem_reasoning
|
100.00%
|
28.08.2026 |
|
SUPERChem
stem_reasoning
|
87.14%
|
28.08.2026 |
|
HorizonMath (pass@4)
stem_reasoning
|
100.00%
|
28.08.2026 |
|
MathArena Apex 2025
stem_reasoning
|
97.34%
|
28.08.2026 |
|
BrokenArXiv
stem_reasoning
|
73.92%
|
28.08.2026 |
|
WorkSpaceBench
general_agent
|
44.05%
|
18.08.2026 |
|
ProgramBench
coding_agent
|
29.03%
|
28.08.2026 |
|
ArXivMath
stem_reasoning
|
100.00%
|
28.08.2026 |