Ask ChatGPT the same question twice and you will often get two different answers, built from different searches, citing different websites. The engines are probabilistic by design. So a single test of "does AI mention my business" is close to worthless: it can show you present when you usually are absent, or absent when you usually are present. In our audits we run every prompt more than once on every engine and report how often you appear, because a rate is something you can act on and retest. A full audit runs every question three times, and no question in anything we deliver is asked fewer than twice. That floor sits in the capture tooling itself, so it is not a promise anyone here can quietly skip. Audits delivered before September 2026 ran a single pass and are labeled that way. The reasoning behind all of it is just arithmetic.
The same question, different answers
Three things vary between runs of an identical prompt. The wording of the answer changes, which matters little. The set of searches the engine runs during fan-out changes, which matters a lot, because a different search pool means different pages get read. And the final citations change, which is the part your customers see. In our audit work it is completely normal to watch a brand get named in one run and skipped in the next, with nothing about the brand or its website having changed in between.
Reasoning depth changes the sources
Not every difference between two answers is chance. ChatGPT's composer carries a Think toggle, and where that toggle sits changes which pages the answer gets built from. Semrush, working with Kevin Indig, ran 100 prompts through GPT-5.2, an older model than the GPT-6 Astra OpenAI announced on September 3, 2026, twice, once with thinking off and once with it on, and found that only 25.6% of the cited domains overlapped for the same prompt. Across those 200 responses the citation rate went from 50% to 68%, and the average number of sources per response went from 2.6 to 4.5.
Three runs cannot answer that. Repeating a prompt tells you how much an engine varies at the setting you used, and it tells you nothing about what the other setting would have returned. So we record the setting rather than assume it. Every ChatGPT run in our audits now stores the state of the Think toggle, read off the composer at the moment the prompt is sent, and the state we capture today is Think off, meaning extended reasoning was not requested. That corresponds to the lower-effort half of the Semrush test.
Reading someone else's AI visibility report right now? Ask them the denominator first. Ours ships with one on every finding.
Get my free auditThe arithmetic of one test
Suppose the truth is that an assistant names you in one out of every three answers about your service. You do not know that yet. You open the app and ask once.
Both outcomes feel like evidence. Neither is. And the failure compounds when you test a change: post a new page, ask once, see your name, and you will credit the page for what may be ordinary variance. Agencies demo this trick live on sales calls, sometimes without knowing it is a trick.
If AI names you one time in three, a single test will tell you whatever you were hoping to hear.
Why three, and not one or ten
Two runs per prompt per engine is the hard floor, enforced in the capture tooling rather than in a policy document, because one run cannot distinguish luck from pattern at all. A full audit pays for a third. Ten runs would be more precise, and for a monthly retainer we do rerun and track over time, but for a diagnosis the marginal precision is not worth tripling the cost. Three runs is the smallest number that separates the outcomes a business owner actually needs to act on:
- 0 of 3. You are effectively invisible for this question. The work starts at finding out whether you were even retrieved.
- 1 or 2 of 3. You are in the rotation. The engine considers you a plausible answer and does not consider you the answer. This is usually the cheapest ground to gain, because the retrieval problem is already solved.
- 3 of 3. You own this question today. The job is defense: keep the page current and watch who is climbing.
The same rule applies to reading anyone else's data, including ours. When a report says "AI recommends your competitor," the first question to ask is: out of how many runs? A finding that does not come with a denominator was a screenshot, and screenshots are how this industry sells fear.
What this looks like in a report
Every finding in an EZ Web audit carries its rate: named in 2 of 3 runs on ChatGPT, 0 of 3 on Perplexity, and so on, per question, per engine. When we later measure whether our work moved anything, the comparison is rate against rate, run the same way. It is slower and roughly three times more expensive for us than screenshotting one lucky answer. It is also the difference between measurement and theater.
Get your appearance rate, not a screenshot
Every question in our free audit is run at least twice per engine, and three times on a full audit. You see the rates, the sources, and the competitors who beat you.
Get my free auditSources
- Google, "AI in Search" (query fan-out, the mechanism behind run-to-run variance in retrieval): blog.google
- Semrush with Kevin Indig, "Only 25% of cited sources overlap between ChatGPT's different reasoning modes [Study]" (100 prompts, 200 responses, GPT-5.2, published June 30, 2026): semrush.com
- The appearance-rate figures in this article are illustrative arithmetic, not survey data. The run-to-run variance itself is directly observable by asking any assistant the same question in fresh sessions.