← leaderboard
GPT-5.4 Mini
SUSPICIOUSOpenAI · openai/gpt-5.4-mini · verdict as of 2026-08-10
Flags
- · report-analysis: score 95.0 vs baseline 100.0 (z=-2.50, drop 5.0%)
Daily test history (90 days, baseline band = trailing mean ± 2σ)
Code Generation
NOT COOKED
Math & Reasoning
NOT COOKED
Structured Extraction
NOT COOKED
Game Design
NOT COOKED
Report Analysis
SUSPICIOUS
Summarization
NOT COOKED
Customer Service
NOT COOKED
Web Design
NOT COOKED
Creative Writing
no data
Instruction Following
NOT COOKED
Serving providers (last 30 days)
striped = multiple providers served this model that day (hover for detail)
Public benchmarks overall 68.2
MMLU-Pro
55.3
GPQA Diamond
88
SWE-bench Verified
—
LMArena Elo
1449
AIME 2025
—
retrieved 2026-07-03 from public sources — see methodology
Recent samples (latest run, one per test case)
No samples from the latest run.