← leaderboard
Kimi K2.6
SUSPICIOUSMoonshot · moonshotai/kimi-k2.6 · verdict as of 2026-07-29
Flags
- · report-analysis: score 83.3 vs baseline 94.7 (z=-1.06, drop 12.0%)
Daily test history (90 days, baseline band = trailing mean ± 2σ)
Structured Extraction
NOT COOKED
Web Design
NOT COOKED
Math & Reasoning
NOT COOKED
Code Generation
NOT COOKED
Instruction Following
NOT COOKED
Report Analysis
SUSPICIOUS
Summarization
NOT COOKED
Game Design
CALIBRATING
Customer Service
NOT COOKED
Creative Writing
CALIBRATING
Serving providers (last 30 days)
striped = multiple providers served this model that day (hover for detail)
Public benchmarks overall 80.4
MMLU-Pro
84.6
GPQA Diamond
90.5
SWE-bench Verified
80.2
LMArena Elo
1461
AIME 2025
—
retrieved 2026-07-03 from public sources — see methodology
Recent samples (latest run, one per test case)
No samples from the latest run.