← leaderboard
Grok 4.3
NOT COOKEDxAI · x-ai/grok-4.3 · verdict as of 2026-08-10
Daily test history (90 days, baseline band = trailing mean ± 2σ)
Code Generation
NOT COOKED
Instruction Following
NOT COOKED
Structured Extraction
NOT COOKED
Report Analysis
NOT COOKED
Summarization
NOT COOKED
Math & Reasoning
NOT COOKED
Customer Service
NOT COOKED
Game Design
NOT COOKED
Web Design
NOT COOKED
Creative Writing
no data
Serving providers (last 30 days)
striped = multiple providers served this model that day (hover for detail)
Public benchmarks overall 70.9
MMLU-Pro
—
GPQA Diamond
87.5
SWE-bench Verified
—
LMArena Elo
1443
AIME 2025
—
retrieved 2026-07-03 from public sources — see methodology
Recent samples (latest run, one per test case)
No samples from the latest run.