Gemini 3.1 Pro

Google

Parameters

—

Architecture

—

Released

—

License

—

Gemini

About

No detailed description available for this model.

Benchmark Scores

Benchmark Score Date
AgentWorldBench MCP
general_agent
27.29%
24.06.2026
AgentWorldBench Search
general_agent
52.08%
24.06.2026
AgentWorldBench Terminal
general_agent
61.79%
24.06.2026
AgentWorldBench SWE
general_agent
69.66%
24.06.2026
AgentWorldBench Web
general_agent
75.79%
24.06.2026
AgentWorldBench OS
general_agent
76.45%
24.06.2026
AgentWorldBench Overall
general_agent
67.12%
24.06.2026
SWE-bench Verified
coding_agent
91.97%
11.06.2026
SWE-bench Pro
coding_agent
67.75%
11.06.2026
Terminal Bench 2.1
coding_agent
77.49%
11.06.2026
SWE Atlas - QnA
coding_agent
7.85%
11.06.2026
NL2Repo-Bench
coding_agent
21.60
11.06.2026
SWE Atlas - TW
coding_agent
39.50%
11.06.2026
SWE-efficiency
coding_agent
17.48%
11.06.2026
LiveSQLBench
coding_agent
84.62%
11.06.2026
CL-bench
coding_agent
57.00%
11.06.2026
VIBE-V2
coding_agent
28.00
11.06.2026
SVG-Bench
coding_agent
69.57%
11.06.2026
PostTrainBench
coding_agent
7.17%
11.06.2026
KernelBench Hard
coding_agent
40.10%
11.06.2026
PaperBench
coding_agent
25.80%
11.06.2026
BrowseComp
general_agent
93.98%
11.06.2026
GDPVal
general_agent
50.33%
11.06.2026
BankerToolBench
general_agent
17.22%
11.06.2026
OfficeQA Pro
general_agent
18.10
11.06.2026
SpreadSheetBench-v1
general_agent
39.01%
11.06.2026
YC-Bench
general_agent
50.00%
11.06.2026
MCP-Atlas
general_agent
77.43%
11.06.2026
Apex-Agents
general_agent
76.80%
11.06.2026
Claw-Eval Avg
coding_agent
69.13%
11.06.2026
OSWorld-Verified
general_agent
82.52%
11.06.2026
OmniDocBench 1.5
document_understanding
96.17%
11.06.2026
MMMU-Pro
vision_language
92.60%
11.06.2026
VideoMMMU
video_understanding
100.00%
11.06.2026
VideoMME (w sub.)
video_understanding
82.93%
11.06.2026
IMO 2025
stem_reasoning
37.92%
11.06.2026
Humanity's Last Exam
stem_reasoning
81.25%
17.06.2026
HLE (with tools)
stem_reasoning
72.25%
17.06.2026
CritPt (no tools)
stem_reasoning
53.94%
17.06.2026
AIME 26
stem_reasoning
98.72%
17.06.2026
HMMT Nov 25
stem_reasoning
91.37%
17.06.2026
HMMT Feb 26
stem_reasoning
86.53%
17.06.2026
IMOAnswerBench
stem_reasoning
82.95%
17.06.2026
GPQA Diamond
stem_reasoning
100.00%
17.06.2026
NL2Repo
coding_agent
35.54%
17.06.2026
DeepSWE 1.1
coding_agent
3.02%
17.06.2026
ProgramBench
coding_agent
52.98%
17.06.2026
Terminal-Bench 2.1 (Terminus-2)
coding_agent
78.17%
17.06.2026
Terminal-Bench 2.1 (Best Reported Harness)
coding_agent
11.81%
17.06.2026
SWE-Marathon
coding_agent
6.12%
17.06.2026
Tool Decathlon
general_agent
76.83%
17.06.2026
DeepSearch QA
general_agent
60.66%
23.08.2026
Toolathlon Verified
general_agent
41.92%
23.08.2026
MCPMark
general_agent
58.13%
23.08.2026
Claw-Eval Pass^3
coding_agent
72.25%
23.08.2026
Terminal-Bench 2.0
coding_agent
100.00%
23.08.2026
SWE-bench Multilingual
coding_agent
85.09%
23.08.2026
SciCode (subtask)
stem_reasoning
100.00%
23.08.2026
OJBench
stem_reasoning
100.00%
23.08.2026
LiveCodeBench v6
stem_reasoning
97.54%
23.08.2026
CharXiv (RQ)
document_understanding
56.13%
23.08.2026
MathVision
vision_language
84.56%
23.08.2026
BabyVision
vision_language
50.65%
23.08.2026
V-Star
vision_language
96.23%
23.08.2026
MMLU-Pro
knowledge
100.00%
26.06.2026
CodeForces
stem_reasoning
87.53%
26.06.2026
Apex-Shortlist (no tools)
stem_reasoning
98.21%
26.06.2026
GDPVal-AA v2
general_agent
71.76%
26.06.2026
USAMO 2026
stem_reasoning
47.58%
11.06.2026
FrontierSWE
coding_agent
17.91%
17.06.2026
AgentWorldBench Android
general_agent
96.36%
24.06.2026

Related Models