How we measure AI visibility.
Every number in an EZ Web audit is labeled with how it was produced and how much weight it can carry. This page is the full protocol behind those labels, published so you can check our work and hold any other tool to the same standard. The instrument that runs it is Citrion, built and operated in-house.
We measure the queries, not just the answers.
When someone asks ChatGPT or Perplexity who to hire, the assistant does not search their words. It writes its own search queries first, a step the industry calls query fan-out, then builds its answer from the handful of pages those queries return. Those queries decide who gets cited, so they are the thing worth measuring.
Citrion records the fan-out queries live, during real sessions, on five of the six engines we audit: ChatGPT, Perplexity, Claude, Grok, and Copilot. Nothing on the observed side is inferred from an API or reconstructed after the fact. On ChatGPT we also read the composer's Think toggle on every run and store its state with the run, because reasoning depth changes which sources an answer cites. Every run we capture today reads Think off, so extended reasoning was not requested. Once the answer renders we also read the model identifier ChatGPT attaches to it and store it with the run exactly as the page wrote it. Our capture account is a free one with no model picker, so when the page names no model the run stores that reason and leaves the model blank. Each run is captured from a stated location with device location sharing off: the run stores the place its IP address resolves to as ipLocation and the sharing state as deviceLocation, so a local answer is read as the one that city gets.
Google does not expose its fan-out. For Google we model the likely queries from the engines we can observe and from Google's own documentation, and every Google number in the report carries a MODELED label. A modeled number is a hypothesis. We never present one as an observation.
The repeatability test.
Ask an AI assistant the same question twice and you get two different answers, built from two different sets of queries. A single run is an anecdote. So every question in an audit runs at least twice, and three times on a full audit. That has been the standard since September 2026; audits delivered before that date ran a single pass per question and say so in their own method notes.
- Each run starts in a fresh chat with no history and no memory of the earlier runs.
- Runs are spaced at least five minutes apart, so we are not replaying a cached result.
- The query sets are compared line by line, run against run.
Queries that appear in every run form the invariant core. This is the engine's stable behavior, and it is the only thing your plan is built on. Queries that appear once or twice are the volatile tail. They stay in the report, labeled as noise, so you can see them without building a plan on them.
Two or three runs is a deliberately small sample. It cannot estimate how often a query appears. It can do something more useful for planning: separate what the engine always does from what it happened to do once.
Why there is no ranking in this report.
The strongest objection to work like ours is that AI answers change on every ask, which would make a single audit worthless. In September 2026 we tested that against our own archive: 38 stored audits, every question and engine pair that had been asked more than once, each run compared against its repeat.
- Whether you appear is reasonably stable. Counting only the pairs where the brand appeared in at least one of the two answers, both answers carried it in 38 of 42 cases when the runs were days apart, and in 73 of 96 cases when the runs came minutes apart in one sitting.
- Which sources the answer cites is not stable. On those same runs the cited domains overlapped by about half within a sitting and by about a third across days. Roughly one comparison in eight shared no sources at all, and about a quarter matched exactly.
- The flattering version of this number is the one we refuse to print. Counting every pair, including the ones where the brand was missing from both answers, the two runs agreed 348 times out of 371. Most of that agreement is absence matching absence, which is easy to get right and tells you nothing. The narrower figure above is the honest one.
That split is the whole reason the report gives you a visibility percentage across repeat runs and never a rank. Position inside an answer sits on the layer that moves; being in the answer at all sits on the layer that holds. Reporting a rank would mean selling you the unstable half as though it were the stable one.
It has a second consequence worth knowing. A finding that you are absent is more durable than a finding that some particular publication cited you, because absence repeated across runs is the most consistent result in the data. We write both up with that difference attached, and every stability figure in a report is the count from your own runs, not a number carried over from this page.
Retrieved is one problem. Cited is another.
Engines fetch more pages than they cite. We log both sets separately, because they fail differently.
If your page was never retrieved, the engine could not find it. That is a coverage problem: crawl access, indexing, or content that misses the queries the engine actually runs.
If your page was retrieved and then passed over, the engine read it and chose other sources. That is a problem on the page itself, and it is usually the faster one to fix.
A report that only says "you were not cited" cannot tell you which of these happened to you. Ours does, for every question in the set.
Named is not the same as recommended.
An assistant can name your business and still talk a buyer out of it in the same sentence. Counting mentions misses that, so every recorded answer that names you is read a second time for how it describes you.
Each of those answers gets one of four labels: favorable, neutral, mixed, or negative. Alongside the label we keep the phrases the answer used to describe you and, when the label is mixed or negative, the drawback it attached. Every quote in the report is checked against the stored answer before it is kept, so the report never prints a sentence an assistant did not say.
The read is done by a language model working from the recorded text, not from a fresh session, and the report says how much of each answer it had to work with. It is a read, not a rate: the labels tell you what a lead hears when they check you with an assistant, and the quotes tell you what to publish in reply.
Every number carries an evidence tier.
Captured live from the engine during a real session. The strongest tier. Fan-out queries, retrieved pages, and citations on the five observed engines all live here.
Our reconstruction where an engine hides its behavior. Google fan-out lives here. Useful for planning, clearly labeled, never dressed up as observed.
Published industry research that shows association without proving cause. It informs recommendations and is always quoted with its limits attached.
Where a sample is too small to support a number, the report says so instead of printing one. When published studies disagree, we quote the range rather than picking the flattering end.
The outside evidence we lean on.
Our recommendations follow the published evidence, including where it is unflattering to common GEO advice.
- Schema markup is not an AI visibility fix. Ahrefs ran a controlled test on 1,885 pages and found that adding structured data moved AI Overviews visibility down 4.6 percent, with no significant change in ChatGPT or AI Mode. We still use schema where it earns its keep, such as rich results and entity resolution. We do not sell it as the thing that gets you cited.
- AI systems read the text on the page. In a searchVIU test, a price that existed only in JSON-LD markup was found by zero of the five AI systems checked. If a fact matters, it goes in the visible copy.
- Visible substance is what lifts citations. In the controlled tests of the Princeton GEO study (Aggarwal et al., KDD 2024), adding quotations, statistics and cited sources to visible text lifted a page's word share in AI answers by 30 to 40 percent. On live Perplexity the best method, quotations, gained 22 percent.
- Mentions track AI visibility more closely than backlinks. Across roughly 75,000 brands, Ahrefs measured a Spearman correlation of 0.664 between web mentions and AI visibility, against 0.218 for backlinks and 0.295 for referring domains. This is correlational, so we treat it as a direction and say so.
- Search rankings track AI mentions far more closely than links do. In January 2025, Seer Interactive ran 10,000 finance and SaaS questions through GPT-4o and compared the brands it mentioned against Google and Bing data. Google organic keyword rankings correlated at 0.65 with being mentioned, Bing organic at 0.56, Google SERP features at 0.41, domain rank at 0.25, and backlinks at 0.10. For service sites the figures ran higher and the order held. This is correlational and covers one model in two industries, so we treat it as a direction and say so.
- Disputed numbers get quoted as ranges. Published estimates of how much AI Overviews overlap with the top ten organic results run from 37 to 54 percent depending on the study. Numbers like this appear in our reports as ranges, never as a single confident figure.
- Some citation share follows licensing deals. Press Ranger and OtterlyAI found licensed publishers taking 48 percent more citations per cited page on ChatGPT, an association in one month of mostly news data. Our own count across 31 stored audits did not show that split, and the detail is in bought vs earned AI citations.
What we do not claim, and what stays private.
Attribution is unsolved. No tool can reliably trace a signed client back to one AI answer, and we do not pretend otherwise. We measure visibility, which is the input you can actually move.
Models change underneath everyone. When an engine ships a new model, its query behavior can shift, so older trend lines get read with caution and a fresh rerun beats an archive.
We do not write or buy reviews, and we do not place links on private blog networks. When we quote a statistic, it links to the study it came from.
And one boundary: this page publishes the protocol, while Citrion, the capture tooling that implements it, stays private. Anyone can follow the recipe by hand. The instrument that makes it repeatable at audit scale is ours.