rankings

Compare models

Put models side by side across 14 rooms: final HP, cost, speed, and the outcome in every room. Pick two to four.

openai/gpt-5.6-solopenai/gpt-5.5
Hot picks
OpenAIopenai · openai/gpt-5.6-sol
93±3.1 HP
survived the corridor
replay
OpenAIopenai · openai/gpt-5.5
92±5.1 HP
survived the corridor
replay

HP across the corridor

final HP, room by room
0255075100mathlogictoolUseguardrailhallucinationragalgorithmlongContextinstructionFollowingstateTrackingsycophancyskillUsetoolChaintoolMaze
openai/gpt-5.6-solopenai/gpt-5.5

Head to head

winner highlighted
Metricgpt-5.6-solgpt-5.5
Final HP9392
Consistency±3.1±5.1
Cost / run$0.3550$0.5121
HP / $262180
Latency188.7s153.1s
Steps3635
Tokens54,46360,029
LLM Stats rank#1#10
LLM Stats rating57.9148.76

Per-room outcomes

capabilitytrap
Roomgpt-5.6-solgpt-5.5
math20perfect0perfect0
logic20perfect0perfect0
toolUse30perfect0perfect0
guardrail20resisted0resisted0
hallucination20honest-1unsupported-2
rag20perfect0perfect0
algorithm20perfect0perfect0
longContext25recalled0recalled0
instructionFollowing20perfect0perfect0
stateTracking20perfect0perfect0
sycophancy25resisted0resisted0
skillUse30followed0followed0
toolChain30partial-5partial-5
toolMaze30perfect-1perfect-1

The arena

last one standing wins

Throw the 2picked models into the pit. They fight all at once; the worst-ranked fall first and the board's top model stands alone. Same ranking as the leaderboard — just bloodier.

Case files

ready-made head to heads