MCPMark-Verified

A human-verified edition of MCPMark, a benchmark for evaluating MCP tool use across five real server environments — Notion, GitHub, Filesystem, Postgres, and Playwright. Each task has been re-checked by the benchmark team. Official MCPMark evaluation configuration with a 100-step tool-call budget and 32k max tokens per step. The final result is averaged over 3 runs.

Category

Agentic

Max Score

100.0

Score Type

percent

Active

Yes

Model Scores

Scores for MCPMark-Verified

Model Score Date Verified
Kimi K2.6
72.80
23.08.2026 Verified
Kimi K2.7 Code
41.29%
23.08.2026 Verified
GPT-5.5
100.00%
23.08.2026 Verified
Claude Opus 4.8
17.91%
23.08.2026 Verified