← leaderboard
MiMo V2.5 Pro
SUSPICIOUSXiaomi · xiaomi/mimo-v2.5-pro · verdict as of 2026-08-10
Flags
- · structured-extraction: score 60.0 vs baseline 83.9 (z=-1.72, drop 28.5%)
- · code-generation: score 80.0 vs baseline 94.7 (z=-1.89, drop 15.5%)
- · instruction-following: score 60.0 vs baseline 87.0 (z=-3.10, drop 31.0%)
Daily test history (90 days, baseline band = trailing mean ± 2σ)
Math & Reasoning
NOT COOKED
Code Generation
SUSPICIOUS
Instruction Following
SUSPICIOUS
Report Analysis
NOT COOKED
Structured Extraction
SUSPICIOUS
Web Design
NOT COOKED
Summarization
NOT COOKED
Customer Service
NOT COOKED
Game Design
NOT COOKED
Creative Writing
no data
Serving providers (last 30 days)
striped = multiple providers served this model that day (hover for detail)
Public benchmarks overall 79.4
MMLU-Pro
84.9
GPQA Diamond
82.6
SWE-bench Verified
78.9
LMArena Elo
1466
AIME 2025
—
retrieved 2026-07-03 from public sources — see methodology
Recent samples (latest run, one per test case)
No samples from the latest run.