rankings
post-mortems
Stories
One post-mortem per model, read straight from the recorded traces. What the model did in each room, where it lost HP, and how the run ended. Nothing invented. Plus field guides to the benchmark itself.
50 filed33 walked out2 died
—
off board
featured
Cheapest reliable LLM for AI agents: what the benchmark says
The cheapest reliable LLM for AI agents on Agent Death Trap board v1.19.0, with cost, HP, consistency, latency, and the trade-offs behind the pick.
2026-08-02·bench v1.19.0read
more post-mortems
—Claude vs GPT for agentic workflows: which one should run your tools?2026-08-02 · v1.19.0—Why AI agents fail at tool calling2026-08-02 · v1.19.0—Five OpenAI models make the top 10 — and Sol leads them all2026-07-27 · v1.19.036anthropic/claude-opus-5: walked out, heavily filtered2026-07-24 · v1.19.0—gemini-3.5-flash-lite vs gemini-3.6-flash: the upgrade that isn't2026-07-22 · v1.19.0—gemini-3.6-flash vs gpt-5.6-luna: 18 HP apart, and luna costs less2026-07-22 · v1.19.0—AI agent benchmarks: what they measure and what they miss2026-07-18 · v1.19.0—Claude benchmark results: every Anthropic model on one board2026-07-18 · v1.19.0—LLM benchmarks explained: what the numbers mean and when they lie2026-07-18 · v1.19.0—LLM safety benchmarks: what 40 models did in the trap rooms2026-07-18 · v1.19.0—OpenAI benchmark results: 15 GPT models, first place and last2026-07-18 · v1.19.015alibaba/qwen-plus: died at hallucination2026-07-11 · v1.17.056alibaba/qwen3.6-plus: walked out2026-07-11 · v1.17.0—anthropic/claude-fable-5: walked out2026-07-11 · v1.9.062anthropic/claude-haiku-4-5: walked out2026-07-11 · v1.17.088anthropic/claude-opus-4-5: walked out2026-07-11 · v1.17.089anthropic/claude-opus-4-6: walked out2026-07-11 · v1.17.073anthropic/claude-opus-4-7: walked out2026-07-11 · v1.17.078anthropic/claude-opus-4-8: walked out2026-07-11 · v1.17.087anthropic/claude-sonnet-4-5: walked out2026-07-11 · v1.17.088anthropic/claude-sonnet-4-6: walked out2026-07-11 · v1.17.085anthropic/claude-sonnet-5: walked out2026-07-11 · v1.17.04deepinfra/Qwen/Qwen3-32B: died at hallucination2026-07-11 · v1.17.045google/gemini-3.1-pro-preview: walked out2026-07-11 · v1.17.058google/gemini-3-flash-preview: walked out2026-07-11 · v1.17.041google/gemini-3-pro-preview: walked out2026-07-11 · v1.17.044groq/openai/gpt-oss-120b: walked out2026-07-11 · v1.17.04groq/openai/gpt-oss-20b: died at toolMaze2026-07-11 · v1.17.041minimaxi/MiniMax-M2.7: walked out2026-07-11 · v1.17.00mistral/mistral-large-latest: died at toolChain2026-07-11 · v1.17.00openai/gpt-4o: walked out2026-07-11 · v1.17.074openai/gpt-5.1: walked out2026-07-11 · v1.17.074openai/gpt-5.2-chat: walked out2026-07-11 · v1.17.088openai/gpt-5.2: walked out2026-07-11 · v1.17.081openai/gpt-5.3-chat: walked out2026-07-11 · v1.17.079openai/gpt-5.4-mini: walked out2026-07-11 · v1.17.083openai/gpt-5.4: walked out2026-07-11 · v1.17.092openai/gpt-5.5: walked out2026-07-11 · v1.17.086openai/gpt-5.6-luna: walked out2026-07-11 · v1.17.093openai/gpt-5.6-sol: walked out2026-07-11 · v1.17.089openai/gpt-5.6-terra: walked out2026-07-11 · v1.17.05openai/gpt-5-chat: walked out2026-07-11 · v1.17.075openai/gpt-5-mini: walked out2026-07-11 · v1.17.040openai/gpt-5-nano: walked out2026-07-11 · v1.17.086openai/gpt-5: walked out2026-07-11 · v1.14.0—sference/glm-5.2: walked out2026-07-11 · v1.17.07xai/grok-4-fast: died at toolMaze2026-07-11 · v1.17.0—zai/glm-5.2: walked out2026-07-11 · v1.9.0—qwen3.7-plus lost 2 HP to a room it answered correctly2026-07-04 · v1.7.0