isitcooked.ai

IS IT COOKED?

The same tests, every model, every day. When a model silently gets worse — quantized, throttled, or “improved” — the scores drop and the verdict flips. Receipts included.

last run 2026-08-10 · nothing cooked today

ModelBench scoreVerdictStructured ExtractionMath & ReasoningCode GenerationInstruction FollowingReport AnalysisSummarizationCustomer ServiceWeb DesignGame DesignCreative Writing
Claude Fable 5Anthropic97.0SUSPICIOUS
Claude Sonnet 5Anthropic91.1NOT COOKED
Gemini 3.1 ProGoogle87.9NOT COOKED
Nex N2 ProNex AGI85.7NOT COOKED
GPT-5.5OpenAI85.2NOT COOKED
GLM 5.2Z.AI82.4NOT COOKED
MiMo V2.5 ProXiaomi79.4SUSPICIOUS
MiniMax M3MiniMax74.6SUSPICIOUS
Grok 4.3xAI70.9NOT COOKED
GPT-5.4 MiniOpenAI68.2SUSPICIOUS
Claude Haiku 4.5Anthropic58.6NOT COOKED
Gemini 3.6 FlashGoogleNOT COOKED55
Grok 4.5xAISUSPICIOUS
Kimi K3Moonshotno data
Muse Spark 1.1MetaNOT COOKED95
GPT-5.6 SolOpenAINOT COOKED
GPT-5.6 TerraOpenAINOT COOKED
GPT-5.6 LunaOpenAINOT COOKED
Claude Opus 5AnthropicSUSPICIOUS9799

Bench score aggregates public benchmarks (LMArena, MMLU-Pro, GPQA, SWE-bench, AIME). Sparklines are 14 days of our own daily tests. Methodology →