Tempe, Arizona
Joshwa Yeong
What I’m actually doing
I got interested in a narrow question: when someone asks ChatGPT for the best option in a category, which companies does it name, and is that stable? It turns out the second half is the hard part. Ask the same question twice and ChatGPT gives you a materially different answer 17% of the time. Gemini does it 28% of the time.
That matters because almost everything being sold as an “AI visibility finding” right now is a single screenshot taken once, which sits comfortably inside that noise. So I built a harness that runs the same buyer-intent question repeatedly across both engines from a fixed location, in neutral sessions, and screenshots every answer at full height with a timestamp.
The rule I run it under only ever deletes findings. It has never manufactured one. In the first wave, six companies went in, three looked like they had something after the first pass, and two survived the second. Most recently I killed a finding on a national retailer after discovering the engines had quietly localised my “neutral” captures to Phoenix — which made an absence look like invisibility when it was the engine being correct.
I’d rather tell you that than show you a case study.
Background
Where this comes from
Marketplace search. Before this I spent my time on Etsy and eRank — ranking products inside a search engine I didn’t control, against a ranking model nobody published, which you can only infer from data you collect yourself. Working out what a black-box ranker rewards by running experiments against it is the same job I’m doing now. The mediator changed. The work didn’t.
Communication, Arizona State University. Which is a less unusual route into measurement work than it sounds, and a more useful one than it looks: most of this job is deciding what a piece of evidence does and does not support.
Building things. The capture harness is Python and Playwright, driving the consumer products rather than the APIs — deliberately, because the whole point is that a prospect can reproduce the result by pasting the prompt into the same product they already use. An API answer is a different system with different retrieval, and quoting one would silently void the finding.
How I work
- Every claim carries a number or a source. If I can’t point at where it came from, I don’t make it.
- Corrections get published, not buried. I revised my own headline instability figure after tripling the sample — it moved from 11%/38% to 17%/28%, and the walk-back is better evidence than the original number was.
- A negative result is a result. The method is worth something precisely because it deletes things.
- No fabricated proof. Not a borrowed logo, not an invented metric, not a client I didn’t have.
Get in touch
Easiest thing: tell me a business and a category and I’ll run a Gap Read on it and send you what comes back. That works whether you’re a prospective client, someone hiring, or just curious what the engines say about you.