All stories

gemini-3.6-flash vs gpt-5.6-luna: 18 HP apart, and luna costs less

2026-07-22 · benchmark v1.19.0

This one is not close. On board v1.19.0, gpt-5.6-luna walked out with 86 HP, rank 11 of 42. gemini-3.6-flash walked out with 68, rank 22. Luna also cost less per run ($0.10 to $0.32), answered faster (122s to 172s), used fewer tokens, and held half the seed spread. There is no axis on the board where the Gemini wins by more than a point.

The record

Measure gpt-5.6-luna gemini-3.6-flash
Final HP 86 ±7.9 68 ±15.3
Rank 11 of 42 22 of 42
Fate walked out walked out
Cost / run $0.1002 $0.3152
HP per dollar 858 216
Latency 122.0s 172.2s
Tokens 59,855 81,986

The trap rooms are even

Start with what the Gemini gets right, because it matters. Both models resisted the guardrail. Both resisted sycophancy. Neither hallucinated: luna graded unsupported (a claim past the document, not an invented fact, 2 HP), the Gemini graded honest (4 HP on average). On the four rooms built to catch models being unsafe or dishonest, these two are peers. Whatever separates them, it is not character.

The gap is execution

The head-to-head page has every room; the gap concentrates in four:

  • Tool maze, the widest split on the page: luna graded perfect and lost 2 HP on average, the Gemini graded partial and lost 13.
  • Math: luna clean, the Gemini averaged 4 HP of damage across seeds despite a perfect representative run.
  • Instruction following: 1 HP against 4, same pattern.
  • Tool chain is the one room that treats them the same: both partial, 5 HP each. Nobody on this board clears that room clean.

For the record, the Gemini's only room wins are logic and long context, one point each.

The seeds tell the same story from another angle. Luna finished at 87 or better on eight of ten seeds, with a floor of 70. The Gemini's median seed is around 67, its floor 45, and its best seed (90) is a day luna has eight times out of ten. Consistency is the difference between a score and a habit.

Which one to run

Luna, and it is not a judgment call on this board: more HP, lower price, lower latency, tighter spread. The case for gemini-3.6-flash is stack alignment (Vertex routing, Google tooling), not numbers. If the Gemini is staying regardless, the cheaper play inside the family is its predecessor: gemini-3.5-flash-lite vs gemini-3.6-flash makes that case.

Luna's full trace is in its post-mortem, written from an earlier board (v1.17.0, 81 HP then; the current board has it at 86). The rest of the OpenAI lineup is in the GPT results roundup, scoring rules in the methodology, all 42 models on the live board.

Current rankings on the leaderboard, scoring details in the methodology.