Kimi Claw 24/7 Bench
In-house benchmark by Moonshot AI for evaluating long-horizon agentic performance in persistent, multi-day coworking tasks. It spans 17 professional scenarios across 610 evaluation points, covering domains such as software engineering, ML research, recruiting, trading, marketing. All tasks are executed through the OpenClaw harness. The final score is the average pass rate across all evaluation points, and is averaged over 3 runs.
Category
Agentic
Max Score
100.0
Score Type
percent
Active
Yes
Model Scores
Scores for Kimi Claw 24/7 Bench
| Model | Score | Date | Verified |
|---|---|---|---|
| Kimi K2.6 |
42.90
|
23.08.2026 | Verified |
| Kimi K2.7 Code |
40.40%
|
23.08.2026 | Verified |
| GPT-5.5 |
100.00%
|
23.08.2026 | Verified |
| Claude Opus 4.8 |
75.76%
|
23.08.2026 | Verified |