rankings

post-mortems

Stories

One post-mortem per model, read straight from the recorded traces. What the model did in each room, where it lost HP, and how the run ended. Nothing invented. Plus field guides to the benchmark itself.

50 filed33 walked out2 died
off board

featured

Cheapest reliable LLM for AI agents: what the benchmark says

The cheapest reliable LLM for AI agents on Agent Death Trap board v1.19.0, with cost, HP, consistency, latency, and the trade-offs behind the pick.

2026-08-02·bench v1.19.0read

more post-mortems

Claude vs GPT for agentic workflows: which one should run your tools?2026-08-02 · v1.19.0Why AI agents fail at tool calling2026-08-02 · v1.19.0Five OpenAI models make the top 10 — and Sol leads them all2026-07-27 · v1.19.036anthropic/claude-opus-5: walked out, heavily filtered2026-07-24 · v1.19.0gemini-3.5-flash-lite vs gemini-3.6-flash: the upgrade that isn't2026-07-22 · v1.19.0gemini-3.6-flash vs gpt-5.6-luna: 18 HP apart, and luna costs less2026-07-22 · v1.19.0AI agent benchmarks: what they measure and what they miss2026-07-18 · v1.19.0Claude benchmark results: every Anthropic model on one board2026-07-18 · v1.19.0LLM benchmarks explained: what the numbers mean and when they lie2026-07-18 · v1.19.0LLM safety benchmarks: what 40 models did in the trap rooms2026-07-18 · v1.19.0OpenAI benchmark results: 15 GPT models, first place and last2026-07-18 · v1.19.015alibaba/qwen-plus: died at hallucination2026-07-11 · v1.17.056alibaba/qwen3.6-plus: walked out2026-07-11 · v1.17.0anthropic/claude-fable-5: walked out2026-07-11 · v1.9.062anthropic/claude-haiku-4-5: walked out2026-07-11 · v1.17.088anthropic/claude-opus-4-5: walked out2026-07-11 · v1.17.089anthropic/claude-opus-4-6: walked out2026-07-11 · v1.17.073anthropic/claude-opus-4-7: walked out2026-07-11 · v1.17.078anthropic/claude-opus-4-8: walked out2026-07-11 · v1.17.087anthropic/claude-sonnet-4-5: walked out2026-07-11 · v1.17.088anthropic/claude-sonnet-4-6: walked out2026-07-11 · v1.17.085anthropic/claude-sonnet-5: walked out2026-07-11 · v1.17.04deepinfra/Qwen/Qwen3-32B: died at hallucination2026-07-11 · v1.17.045google/gemini-3.1-pro-preview: walked out2026-07-11 · v1.17.058google/gemini-3-flash-preview: walked out2026-07-11 · v1.17.041google/gemini-3-pro-preview: walked out2026-07-11 · v1.17.044groq/openai/gpt-oss-120b: walked out2026-07-11 · v1.17.04groq/openai/gpt-oss-20b: died at toolMaze2026-07-11 · v1.17.041minimaxi/MiniMax-M2.7: walked out2026-07-11 · v1.17.00mistral/mistral-large-latest: died at toolChain2026-07-11 · v1.17.00openai/gpt-4o: walked out2026-07-11 · v1.17.074openai/gpt-5.1: walked out2026-07-11 · v1.17.074openai/gpt-5.2-chat: walked out2026-07-11 · v1.17.088openai/gpt-5.2: walked out2026-07-11 · v1.17.081openai/gpt-5.3-chat: walked out2026-07-11 · v1.17.079openai/gpt-5.4-mini: walked out2026-07-11 · v1.17.083openai/gpt-5.4: walked out2026-07-11 · v1.17.092openai/gpt-5.5: walked out2026-07-11 · v1.17.086openai/gpt-5.6-luna: walked out2026-07-11 · v1.17.093openai/gpt-5.6-sol: walked out2026-07-11 · v1.17.089openai/gpt-5.6-terra: walked out2026-07-11 · v1.17.05openai/gpt-5-chat: walked out2026-07-11 · v1.17.075openai/gpt-5-mini: walked out2026-07-11 · v1.17.040openai/gpt-5-nano: walked out2026-07-11 · v1.17.086openai/gpt-5: walked out2026-07-11 · v1.14.0sference/glm-5.2: walked out2026-07-11 · v1.17.07xai/grok-4-fast: died at toolMaze2026-07-11 · v1.17.0zai/glm-5.2: walked out2026-07-11 · v1.9.0qwen3.7-plus lost 2 HP to a room it answered correctly2026-07-04 · v1.7.0