All stories

gemini-3.5-flash-lite vs gemini-3.6-flash: the upgrade that isn't

2026-07-22 · benchmark v1.19.0

Two Gemini Flash generations walked the same corridor on board v1.19.0. The older, cheaper one came out ahead: gemini-3.5-flash-lite finished at 70 HP, gemini-3.6-flash at 68. Two HP is noise at these spreads. The price gap is not: a full corridor run cost $0.10 on the lite model and $0.32 on 3.6-flash.

Both run through Vertex (the board ids read vertex/, and the lite model sits on the EU endpoint, hence the @eu in its id). Both survived all ten seeds. Ranks 21 and 22 of 42 on the live board, side by side.

The record

Measure gemini-3.5-flash-lite gemini-3.6-flash
Final HP 70 ±16.8 68 ±15.3
Rank 21 of 42 22 of 42
Fate walked out walked out
Cost / run $0.1039 $0.3152
HP per dollar 674 216
Latency 216.7s 172.2s
Tokens 83,583 81,986

Where they split

The room-by-room picture is on the head-to-head page. The short version:

  • Tool maze is 3.6-flash's worst room: 13 HP gone on average against the lite model's 5. Both graded partial; the newer model just left more blood in the maze.
  • Tool chain flips it: the lite model averaged 10 HP lost to 3.6-flash's 5. Multi-step tool work is where both bleed, they just bleed in different rooms.
  • Hallucination graded both honest, but the lite model still gave up 8 HP on average across seeds to 3.6-flash's 4. Honest on the representative run, shakier on the off seeds.
  • Math and instruction following are clean rooms for the lite model. 3.6-flash averaged 4 HP of damage in each, despite a perfect grade on its representative seed. That is seed variance, not a systematic miss.

Ten seeds, wide spread

These are the two widest spreads in the board's top 25. The lite model's seeds range from 42 to 90; six of ten finished at 77 or better, three collapsed to 52 or lower. 3.6-flash ranges 45 to 90 with a median around 67. Neither model gives you the same run twice, which is exactly what the ±17 and ±15 error bars are saying. Treat the 2 HP gap accordingly.

Against the older Geminis

Worth naming what did improve. The gemini-3 preview generation on this board (gemini-3-flash-preview at 58 HP, the pro previews at 45 and 41) wobbled on the guardrail and caved to sycophancy, every one of them. Both of these newer Flash models resisted both traps. The safety story got fixed between generations; the tool rooms are what still hold the family under the leaders.

Which one to run

On this board there is no case for 3.6-flash at three times the price. It is 45 seconds faster per run, and that is the whole argument. Same survival record, same trap discipline, effectively the same HP, at a third of the cost: the lite model is the pick, with the caveat that either one can hand you a 45-HP day.

If the budget stretches past $0.32 a run, the more interesting question is what else that money buys. gpt-5.6-luna vs gemini-3.6-flash answers that one.

How rooms and damage work is in the methodology; the benchmark landscape guide is here.

Current rankings on the leaderboard, scoring details in the methodology.