proses

Writing · correction

We measured Gemini as 3× less stable than ChatGPT. Then we tripled the sample.

Our first per-engine instability figure said ChatGPT flipped 11% of answers on a repeat ask and Gemini 38% — about 3.5× worse. At roughly triple the sample it came back 17% and 28%, about 1.6×. The direction held. The magnitude more than halved, and one conclusion we had drawn from it is now dead.

What the number is

Ask an answer engine the same buyer-intent question twice and you do not reliably get the same answer. Different companies get named, in a different order, sometimes with a different recommendation. We wanted to know how often, because everything else we do depends on it: if a brand is absent from one run, that is only a finding if absence is more than noise.

So we measure it directly. A cell is one question on one engine. If the set of brands named changes materially between repeat runs of that cell, the cell flipped. The instability figure is the share of cells that flip.

The first number, and why it was wrong

At 27 ChatGPT cells and 26 Gemini cells we got 11% and 38%. Gemini looked about three and a half times less stable. That is a striking result and it is the kind of thing that ends up in a headline.

It was computed on a sample small enough that a handful of cells moved it several points, and we knew that at the time, and we quoted it anyway.

The second number

At 41 ChatGPT cells and 40 Gemini cells — roughly a 50% larger sample on each engine, after a second pass landed — it came back 17% and 28%.

Instability figures before and after the larger sample
First read Current
ChatGPT cells that flip11%17%
Gemini cells that flip38%28%
Sample (ChatGPT / Gemini cells)27 / 2641 / 40
Ratio between the enginesabout 3.5×about 1.6×

The 11% / 38% figures above are superseded. They appear here because this is the correction; they are wrong anywhere else.

What the correction killed

This is the part that matters, and it is the part a quiet edit would have hidden.

On the first number we had written that Gemini is the binding constraint on nearly every disqualification — that when a finding died, it was almost always Gemini that killed it, because Gemini was so much noisier. At 38% against 11% that was a reasonable read.

At 28% against 17% it is not supported. The engines are much closer than we said, and the operative fact turns out to be a different one: ChatGPT is not a reliable anchor either. Roughly one ChatGPT cell in six moves. We had been treating it as the stable reference point against which Gemini looked erratic, and it isn't.

So the headline number moved, and a conclusion we had built on top of it was deleted. That second thing is the real cost of the error, and it is invisible if you just swap the digits.

What it didn’t change

The rule. Before any of this was measured, the rule was already: a single run is a candidate, not a finding, and nothing gets mechanism language until it holds across three or more repetitions on both engines.

That rule was never justified by the number. It was justified empirically, and expensively. In our first wave we tested six companies. After one pass, three looked like they carried a finding. After the confirming pass, two did. We have watched a confident read get overturned inside an hour, more than once. The percentages describe why that keeps happening; they are not what makes the rule correct.

Why publish it

Partly because we said we would. Every claim carries a number or a source, and a number that moves has to move in public or the first commitment was decorative.

But mostly because of what it demonstrates. A measurement that has never been revised is not obviously trustworthy — it might just be one nobody has stress-tested. A measurement that got publicly corrected upward in sample size, with the conclusion it invalidated named out loud, tells you considerably more about how it was produced.

The walk-back is better evidence than the original number ever was. That is not a consolation. It is the actual reason it is on the website.

Both figures are recomputed from our capture corpus: 514 captures, 511 successful, 14 companies, two engines — ChatGPT and Gemini Flash, the free default tier. Those are the only two engines we have run and the only two we claim.

Related

Same mechanism, different mediator

What a Gap Read is — including a finding we withdrew.