All stories

OpenAI benchmark results: 15 GPT models, first place and last

2026-07-18 · benchmark v1.19.0

OpenAI owns both ends of the current board. gpt-5.6-sol leads all 40 models with 93 HP, and gpt-4o sits at the bottom with 0, dead in the sycophancy room. No other vendor spans first place and last.

This page is the OpenAI benchmark record on Agent Death Trap board v1.19.0 (40 models, 10 seeds each, run 2026-07-17): every GPT model's rank, HP, and the pattern that emerges when you line up fifteen of them.

The GPT board, v1.19.0

Rank Model Final HP Fate
1 gpt-5.6-sol 93 ±3.1 walked out
2 gpt-5.5 92 ±5.1 walked out
5 gpt-5.6-terra 89 ±7.2 walked out
8 gpt-5.2 88 ±3.0 walked out
10 gpt-5 86 ±4.7 walked out
11 gpt-5.6-luna 86 ±7.9 walked out
13 gpt-5.4 83 ±8.8 walked out
14 gpt-5.3-chat 81 ±6.9 walked out
15 gpt-5.4-mini 79 ±10.5 walked out
17 gpt-5-mini 75 ±5.3 walked out
18 gpt-5.1 74 ±5.4 walked out
19 gpt-5.2-chat 74 ±11.6 walked out
29 gpt-5-nano 40 ±9.5 walked out
35 gpt-5-chat 5 ±13.8 died at toolMaze
40 gpt-4o 0 ±0 died at sycophancy

The gpt-5.6 triplet

OpenAI shipped three 5.6 variants and the corridor separates them cleanly. Sol wins the whole board at 93 and is nearly as steady as it is strong (±3.1). Terra lands at 89 and luna at 86, but both swing much harder between seeds (±7.2 and ±7.9). Same family, same rooms, different temperament: sol repeats its run, terra and luna have good days and bad days. On a GPT-5 benchmark chart with a single number, that difference is invisible.

Chat variants pay for it

Line up each base model against its chat build and the pattern is hard to miss:

  • gpt-5.2 scores 88; gpt-5.2-chat scores 74.
  • gpt-5 scores 86; gpt-5-chat scores 5 and died in the tool maze.

The chat tunings trade agent discipline for conversational polish, and this corridor prices that trade. If you are picking a model to hand tools to, the base line is the one doing the work.

Where the GPT line is strong, and where it cracks

Sycophancy is a solved problem for the modern line. Fourteen of fifteen GPT models resisted the push-back trap. The one that caved is gpt-4o, and it died there.

The guardrail wobbles more than it should. gpt-5, gpt-5.1, and gpt-5-mini all wobbled under pressure, and gpt-5-nano got fully manipulated. Not fatal at the top of the line, but it is the recurring dent in otherwise clean runs.

Hallucination splits by generation. Eight of the fifteen graded honest. The two that invented facts outright are the two oldest designs on the list, gpt-5-chat and gpt-4o. Both are also the two that died.

gpt-4o deserves its own sentence. It failed three traps in one run: leaked at the guardrail, hallucinated, then caved to a confidently wrong user and bled out. Zero HP, ±0 across all ten seeds. It is the only model on the board killed by a trap room rather than by tool work, and a clean reminder that a model can feel fluent in chat and still be the most manipulable thing on the board.

GPT vs the field

The natural head-to-heads, room by room:

Every model above links to its run page, and each has a trace-based post-mortem in the stories archive. The live leaderboard has all 40 models, the methodology explains the rooms and the rubric, and the Anthropic side of the story is in the Claude benchmark results.

Current rankings on the leaderboard, scoring details in the methodology.