Writing · correction
We measured Gemini as 3× less stable than ChatGPT. Then we tripled the sample.
What the number is
Ask an answer engine the same buyer-intent question twice and you do not reliably get the same answer. Different companies get named, in a different order, sometimes with a different recommendation. We wanted to know how often, because everything else we do depends on it: if a brand is absent from one run, that is only a finding if absence is more than noise.
So we measure it directly. A cell is one question on one engine. If the set of brands named changes materially between repeat runs of that cell, the cell flipped. The instability figure is the share of cells that flip.
The first number, and why it was wrong
At 27 ChatGPT cells and 26 Gemini cells we got 11% and 38%. Gemini looked about three and a half times less stable. That is a striking result and it is the kind of thing that ends up in a headline.
It was computed on a sample small enough that a handful of cells moved it several points, and we knew that at the time, and we quoted it anyway.
The second number
At 41 ChatGPT cells and 40 Gemini cells — roughly a 50% larger sample on each engine, after a second pass landed — it came back 17% and 28%.
| First read | Current | |
|---|---|---|
| ChatGPT cells that flip | 11% | 17% |
| Gemini cells that flip | 38% | 28% |
| Sample (ChatGPT / Gemini cells) | 27 / 26 | 41 / 40 |
| Ratio between the engines | about 3.5× | about 1.6× |
The 11% / 38% figures above are superseded. They appear here because this is the correction; they are wrong anywhere else.
What the correction killed
This is the part that matters, and it is the part a quiet edit would have hidden.
On the first number we had written that Gemini is the binding constraint on nearly every disqualification — that when a finding died, it was almost always Gemini that killed it, because Gemini was so much noisier. At 38% against 11% that was a reasonable read.
At 28% against 17% it is not supported. The engines are much closer than we said, and the operative fact turns out to be a different one: ChatGPT is not a reliable anchor either. Roughly one ChatGPT cell in six moves. We had been treating it as the stable reference point against which Gemini looked erratic, and it isn't.
So the headline number moved, and a conclusion we had built on top of it was deleted. That second thing is the real cost of the error, and it is invisible if you just swap the digits.
What it didn’t change
The rule. Before any of this was measured, the rule was already: a single run is a candidate, not a finding, and nothing gets mechanism language until it holds across three or more repetitions on both engines.
That rule was never justified by the number. It was justified empirically, and expensively. In our first wave we tested six companies. After one pass, three looked like they carried a finding. After the confirming pass, two did. We have watched a confident read get overturned inside an hour, more than once. The percentages describe why that keeps happening; they are not what makes the rule correct.
Why publish it
Partly because we said we would. Every claim carries a number or a source, and a number that moves has to move in public or the first commitment was decorative.
But mostly because of what it demonstrates. A measurement that has never been revised is not obviously trustworthy — it might just be one nobody has stress-tested. A measurement that got publicly corrected upward in sample size, with the conclusion it invalidated named out loud, tells you considerably more about how it was produced.
The walk-back is better evidence than the original number ever was. That is not a consolation. It is the actual reason it is on the website.
Both figures are recomputed from our capture corpus: 514 captures, 511 successful, 14 companies, two engines — ChatGPT and Gemini Flash, the free default tier. Those are the only two engines we have run and the only two we claim.