Kimi Claw 24/7 Bench

In-house benchmark by Moonshot AI for evaluating long-horizon agentic performance in persistent, multi-day coworking tasks. It spans 17 professional scenarios across 610 evaluation points, covering domains such as software engineering, ML research, recruiting, trading, marketing. All tasks are executed through the OpenClaw harness. The final score is the average pass rate across all evaluation points, and is averaged over 3 runs.

Category

Agentic

Max Score

100.0

Score Type

percent

Active

Yes

Model Scores

Scores for Kimi Claw 24/7 Bench

Model Score Date Verified
Kimi K2.6
42.90
23.08.2026 Verified
Kimi K2.7 Code
40.40%
23.08.2026 Verified
GPT-5.5
100.00%
23.08.2026 Verified
Claude Opus 4.8
75.76%
23.08.2026 Verified