OpenAI owns both ends of the current board. gpt-5.6-sol leads all 40 models with 93 HP, and gpt-4o sits at the bottom with 0, dead in the sycophancy room. No other vendor spans first place and last.
This page is the OpenAI benchmark record on Agent Death Trap board v1.19.0 (40 models, 10 seeds each, run 2026-07-17): every GPT model's rank, HP, and the pattern that emerges when you line up fifteen of them.
The GPT board, v1.19.0
| Rank | Model | Final HP | Fate |
|---|---|---|---|
| 1 | gpt-5.6-sol | 93 ±3.1 | walked out |
| 2 | gpt-5.5 | 92 ±5.1 | walked out |
| 5 | gpt-5.6-terra | 89 ±7.2 | walked out |
| 8 | gpt-5.2 | 88 ±3.0 | walked out |
| 10 | gpt-5 | 86 ±4.7 | walked out |
| 11 | gpt-5.6-luna | 86 ±7.9 | walked out |
| 13 | gpt-5.4 | 83 ±8.8 | walked out |
| 14 | gpt-5.3-chat | 81 ±6.9 | walked out |
| 15 | gpt-5.4-mini | 79 ±10.5 | walked out |
| 17 | gpt-5-mini | 75 ±5.3 | walked out |
| 18 | gpt-5.1 | 74 ±5.4 | walked out |
| 19 | gpt-5.2-chat | 74 ±11.6 | walked out |
| 29 | gpt-5-nano | 40 ±9.5 | walked out |
| 35 | gpt-5-chat | 5 ±13.8 | died at toolMaze |
| 40 | gpt-4o | 0 ±0 | died at sycophancy |
The gpt-5.6 triplet
OpenAI shipped three 5.6 variants and the corridor separates them cleanly. Sol wins the whole board at 93 and is nearly as steady as it is strong (±3.1). Terra lands at 89 and luna at 86, but both swing much harder between seeds (±7.2 and ±7.9). Same family, same rooms, different temperament: sol repeats its run, terra and luna have good days and bad days. On a GPT-5 benchmark chart with a single number, that difference is invisible.
Chat variants pay for it
Line up each base model against its chat build and the pattern is hard to miss:
- gpt-5.2 scores 88; gpt-5.2-chat scores 74.
- gpt-5 scores 86; gpt-5-chat scores 5 and died in the tool maze.
The chat tunings trade agent discipline for conversational polish, and this corridor prices that trade. If you are picking a model to hand tools to, the base line is the one doing the work.
Where the GPT line is strong, and where it cracks
Sycophancy is a solved problem for the modern line. Fourteen of fifteen GPT models resisted the push-back trap. The one that caved is gpt-4o, and it died there.
The guardrail wobbles more than it should. gpt-5, gpt-5.1, and gpt-5-mini all wobbled under pressure, and gpt-5-nano got fully manipulated. Not fatal at the top of the line, but it is the recurring dent in otherwise clean runs.
Hallucination splits by generation. Eight of the fifteen graded honest. The two that invented facts outright are the two oldest designs on the list, gpt-5-chat and gpt-4o. Both are also the two that died.
gpt-4o deserves its own sentence. It failed three traps in one run: leaked at the guardrail, hallucinated, then caved to a confidently wrong user and bled out. Zero HP, ±0 across all ten seeds. It is the only model on the board killed by a trap room rather than by tool work, and a clean reminder that a model can feel fluent in chat and still be the most manipulable thing on the board.
GPT vs the field
The natural head-to-heads, room by room:
- gpt-5.6-sol vs gpt-5.5, first place against second
- gpt-5.6-sol vs claude-opus-4-6, the board leader against the best Claude
- claude-sonnet-4-6 vs gpt-5.2, the mid-table rivalry
Every model above links to its run page, and each has a trace-based post-mortem in the stories archive. The live leaderboard has all 40 models, the methodology explains the rooms and the rubric, and the Anthropic side of the story is in the Claude benchmark results.