Eight Claude models walked the corridor on board v1.19.0. All eight survived. None of them won.
This page is the Claude benchmark record in one place: where each Anthropic model ranks, what its trap-room record looks like, and where it actually loses HP. Numbers come from the published board (40 models, 10 seeds each, run 2026-07-17); each model links to its full run page and its post-mortem.
The Claude board, v1.19.0
| Rank | Model | Final HP | Fate |
|---|---|---|---|
| 3 | claude-opus-4-6 | 89 ±2.1 | walked out |
| 6 | claude-opus-4-5 | 88 ±2.4 | walked out |
| 7 | claude-sonnet-4-6 | 88 ±3.3 | walked out |
| 9 | claude-sonnet-4-5 | 87 ±2.9 | walked out |
| 12 | claude-sonnet-5 | 85 ±5.5 | walked out |
| 16 | claude-opus-4-8 | 78 ±6.9 | walked out |
| 20 | claude-opus-4-7 | 73 ±4.7 | walked out |
| 21 | claude-haiku-4-5 | 62 ±11.2 | walked out |
The best Claude on the board is claude-opus-4-6 at 89 HP, rank 3 overall, behind gpt-5.6-sol (93) and gpt-5.5 (92). Its ±2.1 spread across 10 seeds is the steadiest run in the top 10: it does not have great days and bad days, it just does the same thing every time.
The trap record is the story
Claude's signature on this board is trap discipline:
- Guardrail: all 8 resisted. No Claude model leaked under pressure, on a board where 21 of 40 models lost HP in that room.
- Sycophancy: none caved. Seven resisted clean; haiku-4-5 wobbled. Eleven models on the board fully caved, including every Gemini.
- Hallucination: zero hallucinated outcomes. Six graded honest; sonnet-5 and opus-4-8 graded unsupported, which means a claim that went beyond the document, not an invented fact.
- Long context: seven recalled clean; haiku-4-5 went partial.
Across 32 trap-room gradings (8 models, 4 traps), not one full failure. No other vendor with more than two models on the board can say that. The mechanics of these rooms are covered in the safety benchmark guide.
So where does the HP go?
Tool work. Like every model on the board, the Claudes bleed in the multi-step tool chain (nobody cleared it perfectly, 40 of 40) and drop a few points in the tool maze. Opus-4-7 and opus-4-8 also gave up points on instruction following, which is what separates them from the 88-89 HP cluster.
Two honest observations from inside the family:
- Newer is not automatically better. Opus-4-8 (78) sits eleven points under opus-4-6 (89) on this corridor, and opus-4-7 (73) sits under both. The rooms punish discipline slips, and the discipline profile changes between versions.
- Haiku is the budget pick with a real gap. At 62 HP it survives comfortably, but it is the only Claude that wobbles under social pressure and loses thread in long context. Fine for cheap capability work; think twice before handing it secrets.
One footnote: claude-fable-5 ran on an earlier board (v1.9.0) and refused its way through five rooms, ending at 20 HP. It is not on the current board. The full story is in its post-mortem.
Claude vs the field
Head-to-head pages exist for every top-20 pairing, with room-by-room outcomes side by side:
- gpt-5.6-sol vs claude-opus-4-6, first place against the best Claude
- gpt-5.5 vs claude-opus-4-6
- claude-opus-4-6 vs kimi-k3
Per-model post-mortems, written from the recorded traces: opus-4-6, opus-4-5, sonnet-4-6, sonnet-4-5, sonnet-5, opus-4-8, opus-4-7, haiku-4-5.
The live leaderboard has all 40 models; how the scoring works is in the methodology. For the other side of the rivalry, see the OpenAI GPT results.