This one is not close. On board v1.19.0, gpt-5.6-luna walked out with 86 HP, rank 11 of 42. gemini-3.6-flash walked out with 68, rank 22. Luna also cost less per run ($0.10 to $0.32), answered faster (122s to 172s), used fewer tokens, and held half the seed spread. There is no axis on the board where the Gemini wins by more than a point.
The record
| Measure | gpt-5.6-luna | gemini-3.6-flash |
|---|---|---|
| Final HP | 86 ±7.9 | 68 ±15.3 |
| Rank | 11 of 42 | 22 of 42 |
| Fate | walked out | walked out |
| Cost / run | $0.1002 | $0.3152 |
| HP per dollar | 858 | 216 |
| Latency | 122.0s | 172.2s |
| Tokens | 59,855 | 81,986 |
The trap rooms are even
Start with what the Gemini gets right, because it matters. Both models resisted the guardrail. Both resisted sycophancy. Neither hallucinated: luna graded unsupported (a claim past the document, not an invented fact, 2 HP), the Gemini graded honest (4 HP on average). On the four rooms built to catch models being unsafe or dishonest, these two are peers. Whatever separates them, it is not character.
The gap is execution
The head-to-head page has every room; the gap concentrates in four:
- Tool maze, the widest split on the page: luna graded perfect and lost 2 HP on average, the Gemini graded partial and lost 13.
- Math: luna clean, the Gemini averaged 4 HP of damage across seeds despite a perfect representative run.
- Instruction following: 1 HP against 4, same pattern.
- Tool chain is the one room that treats them the same: both partial, 5 HP each. Nobody on this board clears that room clean.
For the record, the Gemini's only room wins are logic and long context, one point each.
The seeds tell the same story from another angle. Luna finished at 87 or better on eight of ten seeds, with a floor of 70. The Gemini's median seed is around 67, its floor 45, and its best seed (90) is a day luna has eight times out of ten. Consistency is the difference between a score and a habit.
Which one to run
Luna, and it is not a judgment call on this board: more HP, lower price, lower latency, tighter spread. The case for gemini-3.6-flash is stack alignment (Vertex routing, Google tooling), not numbers. If the Gemini is staying regardless, the cheaper play inside the family is its predecessor: gemini-3.5-flash-lite vs gemini-3.6-flash makes that case.
Luna's full trace is in its post-mortem, written from an earlier board (v1.17.0, 81 HP then; the current board has it at 86). The rest of the OpenAI lineup is in the GPT results roundup, scoring rules in the methodology, all 42 models on the live board.