← leaderboard
Grok 4.5
SUSPICIOUSxAI · x-ai/grok-4.5 · verdict as of 2026-08-10
Flags
- · code-generation: score 80.0 vs baseline 92.0 (z=-1.54, drop 13.0%)
Daily test history (90 days, baseline band = trailing mean ± 2σ)
Report Analysis
NOT COOKED
Summarization
NOT COOKED
Code Generation
SUSPICIOUS
Math & Reasoning
NOT COOKED
Structured Extraction
NOT COOKED
Customer Service
NOT COOKED
Instruction Following
NOT COOKED
Game Design
NOT COOKED
Web Design
NOT COOKED
Creative Writing
no data
Serving providers (last 30 days)
striped = multiple providers served this model that day (hover for detail)
Recent samples (latest run, one per test case)
No samples from the latest run.