← leaderboard
GPT-5.5
NOT COOKEDOpenAI · openai/gpt-5.5 · verdict as of 2026-08-10
Daily test history (90 days, baseline band = trailing mean ± 2σ)
Instruction Following
NOT COOKED
Report Analysis
NOT COOKED
Summarization
NOT COOKED
Game Design
no data
Math & Reasoning
NOT COOKED
Code Generation
NOT COOKED
Customer Service
NOT COOKED
Web Design
NOT COOKED
Creative Writing
no data
Structured Extraction
NOT COOKED
Serving providers (last 30 days)
striped = multiple providers served this model that day (hover for detail)
Public benchmarks overall 85.2
MMLU-Pro
—
GPQA Diamond
93.6
SWE-bench Verified
—
LMArena Elo
1475
AIME 2025
—
retrieved 2026-07-03 from public sources — see methodology
Recent samples (latest run, one per test case)
No samples from the latest run.