All stories

Five OpenAI models make the top 10 — and Sol leads them all

2026-07-27 · benchmark v1.19.0

Five OpenAI models make the top 10 — and Sol leads them all

OpenAI did not just take first place on the current Agent Death Trap board. It put five models in the top 10.

gpt-5.6-sol leads with 93 HP, one point ahead of gpt-5.5. Three models share 89 HP behind them: Anthropic's Claude Opus 4.6, Moonshot's Kimi K3, and OpenAI's gpt-5.6-terra.

The headline is OpenAI's depth. The more interesting detail is how little separates the entire group.

The top 10

Rank Model Final HP
1 openai/gpt-5.6-sol 93
2 openai/gpt-5.5 92
3 anthropic/claude-opus-4-6 89
4 moonshot/kimi-k3 89
5 openai/gpt-5.6-terra 89
6 anthropic/claude-opus-4-5 88
7 anthropic/claude-sonnet-4-6 88
8 openai/gpt-5.2 88
9 anthropic/claude-sonnet-4-5 87
10 openai/gpt-5 86

Five are from OpenAI, four are from Anthropic, and one is from Moonshot. No other provider appears in the top 10.

A narrow lead, not a runaway

Only seven HP separate first place from tenth. Sol's 93 is the best result on the board, but it is not an uncontested tier of its own: gpt-5.5 sits one point behind, while the next six models are packed into the 88–89 range.

That compression matters. A leaderboard can make adjacent ranks look more decisive than they are. Here, third through fifth are tied on final HP, and sixth through eighth are tied one point lower. The rank gives the order; the station outcomes explain the models.

OpenAI wins on depth

OpenAI holds ranks 1, 2, 5, 8, and 10. Its representation stretches from the winner to the bottom edge of the top 10, with every listed model finishing at 86 HP or better.

Anthropic is nearly as present. Claude models occupy four places, including two Opus and two Sonnet variants. Claude Opus 4.6 is the highest-ranked non-OpenAI model at 89 HP.

Kimi K3 is the outlier in the provider split. It is Moonshot's only entry in this group, but its 89 HP puts it level with Claude Opus 4.6 and gpt-5.6-terra.

What “survival” means here

Every model enters the corridor with 100 HP and faces the same 14 deterministic stations. The rooms test capability, safety and alignment, and robustness. A station returns a fixed outcome; the benchmark rubric converts that outcome into damage.

That makes final HP a compact summary, not the whole diagnosis. Two models can finish with the same score after taking damage in different rooms for different reasons. The useful question is not only who lost less HP? It is also where did they lose it, and why?

The live ranking shows the full board. You can open any model above for its station-by-station run, compare the leaders directly in the comparison view, or read the complete rules in the methodology.

Current rankings on the leaderboard, scoring details in the methodology.

Get told when a new model runs

One email when a new benchmark version lands or a new model walks the corridor. Nothing else.