Methodology

How we measure AI visibility.

Every number in an EZ Web audit is labeled with how it was produced and how much weight it can carry. This page is the full protocol behind those labels, published so you can check our work and hold any other tool to the same standard.

Step one

We measure the queries, not just the answers.

When someone asks ChatGPT or Perplexity who to hire, the assistant does not search their words. It writes its own search queries first, a step the industry calls query fan-out, then builds its answer from the handful of pages those queries return. Those queries decide who gets cited, so they are the thing worth measuring.

We record the fan-out queries live, during real sessions, on five engines: ChatGPT, Perplexity, Claude, Grok, and Copilot. Nothing on the observed side is inferred from an API or reconstructed after the fact.

Google does not expose its fan-out. For Google we model the likely queries from the engines we can observe and from Google's own documentation, and every Google number in the report carries a MODELED label. A modeled number is a hypothesis. We never present one as an observation.

Step two

The repeatability test.

Ask an AI assistant the same question twice and you get two different answers, built from two different sets of queries. A single run is an anecdote. So every question in an audit runs three times:

  • Each run starts in a fresh chat with no history and no memory of the earlier runs.
  • Runs are spaced at least five minutes apart, so we are not replaying a cached result.
  • The three query sets are compared line by line.

Queries that appear in all three runs form the invariant core. This is the engine's stable behavior, and it is the only thing your plan is built on. Queries that appear once or twice are the volatile tail. They stay in the report, labeled as noise, so you can see them without building a plan on them.

Three runs is a deliberately small sample. It cannot estimate how often a query appears. It can do something more useful for planning: separate what the engine always does from what it happened to do once.

Step three

Retrieved is one problem. Cited is another.

Engines fetch more pages than they cite. We log both sets separately, because they fail differently.

If your page was never retrieved, the engine could not find it. That is a coverage problem: crawl access, indexing, or content that misses the queries the engine actually runs.

If your page was retrieved and then passed over, the engine read it and chose other sources. That is a problem on the page itself, and it is usually the faster one to fix.

A report that only says "you were not cited" cannot tell you which of these happened to you. Ours does, for every question in the set.

Step four

Every number carries an evidence tier.

Observed

Captured live from the engine during a real session. The strongest tier. Fan-out queries, retrieved pages, and citations on the five observed engines all live here.

Modeled

Our reconstruction where an engine hides its behavior. Google fan-out lives here. Useful for planning, clearly labeled, never dressed up as observed.

Correlational

Published industry research that shows association without proving cause. It informs recommendations and is always quoted with its limits attached.

Where a sample is too small to support a number, the report says so instead of printing one. When published studies disagree, we quote the range rather than picking the flattering end.

The research

The outside evidence we lean on.

Our recommendations follow the published evidence, including where it is unflattering to common GEO advice.

  • Schema markup is not an AI visibility fix. Ahrefs ran a controlled test on 1,885 pages and found that adding structured data moved AI Overviews visibility down 4.6 percent, with no significant change in ChatGPT or AI Mode. We still use schema where it earns its keep, such as rich results and entity resolution. We do not sell it as the thing that gets you cited.
  • AI systems read the text on the page. In a searchVIU test, a price that existed only in JSON-LD markup was found by zero of the five AI systems checked. If a fact matters, it goes in the visible copy.
  • Visible substance is what lifts citations. The Princeton GEO study (Aggarwal et al., KDD 2024) measured visibility lifts of 22 to 41 percent from adding quotations, statistics, and cited sources to the visible text, depending on the tactic and the query category.
  • Mentions track AI visibility more closely than backlinks. Across roughly 75,000 brands, Ahrefs measured a Spearman correlation of 0.664 between web mentions and AI visibility, against 0.218 for referring domains. This is correlational, so we treat it as a direction and say so.
  • Disputed numbers get quoted as ranges. Published estimates of how much AI Overviews overlap with the top ten organic results run from 37 to 54 percent depending on the study. Numbers like this appear in our reports as ranges, never as a single confident figure.
The fine print

What we do not claim, and what stays private.

Attribution is unsolved. No tool can reliably trace a signed client back to one AI answer, and we do not pretend otherwise. We measure visibility, which is the input you can actually move.

Models change underneath everyone. When an engine ships a new model, its query behavior can shift, so older trend lines get read with caution and a fresh rerun beats an archive.

And one boundary: this page publishes the protocol, while the capture tooling that implements it stays private. Anyone can follow the recipe by hand. The harness that makes it repeatable at audit scale is ours.

See it on your site

Want this protocol run on your business?