How to Measure AI Visibility: The Metrics That Actually Matter
Run the same AI visibility scan against the same homepage twice in one week, with nothing on the site changed in between, and the two reports can disagree — not by a rounding error, but by several points. The site did not get better or worse in that week. Part of what got measured did.
One number, at least three different measurements
A visibility score usually folds together things that behave nothing alike. Whether an AI crawler can reach a page is a fact — it’s either allowed or blocked, and it doesn’t shift on its own between two scans an hour apart. Whether structured data validates is the same kind of fact. Whether a model cites or recommends the business when asked, on the other hand, comes from a live, open-ended question put to a language model, and a model doesn’t return an identical answer to an open-ended prompt every time it’s asked.
Treating all three as one trend line hides which part actually moved. A score that drops five points could mean a crawler rule broke, or it could mean a model answered a recommendation question differently this week than last — and those call for completely different fixes.
A zero is not always a zero
Underneath a visibility report, each individual check an audit engine runs settles into one of a few outcomes: it found something, it found nothing, or it never got an answer at all. The middle one and the last one look identical on a dashboard if nobody kept them apart — both can render as a blank field or a zero — but they mean opposite things. “Found nothing” is an authoritative answer: the source was asked and it genuinely holds no data for that question, for example a domain with no entry in a public registry. “Never got an answer” means a request timed out, a service the check depends on returned a server error, or the run ran out of time before it got there. Nothing was learned either way, and reporting that as a confirmed zero is a different kind of wrong than reporting a stable fact.
A well-built report keeps these two apart on purpose, because collapsing them tells a reader “checked, and clean” when the honest answer is “not checked.” A business reading a report that says “no structured data found” should be able to tell whether that’s a real gap to fix or a source that simply didn’t answer this time.
Basic and deep scans aren’t measuring the same slice of the checklist
The same principle scales up to the whole score, not just one field in it. A basic scan and a full deep scan don’t run the same set of checks — a basic run skips the crawl, the browser render and most off-site sources — so comparing a basic score against a deep score of the same site, or against a different site’s deep score, isn’t comparing like with like. What a GEO audit actually checks covers how that gap gets reflected in the number itself; the short version here is that the tier a score came from matters as much as the score.
Citation and recommendation need more than one question
The model-facing half of a visibility score usually splits into two separate things: whether a model describes the business accurately when asked about it directly, and whether the model brings the business up unprompted when asked a category question that never names it. Because both come from asking a live model something open-ended, one answer is a sample of one. A closer look at that split goes into why a business can pass the first test and fail the second — the point that matters here is that either number needs the same questions repeated, phrased a few different ways, across more than one model, before it’s safe to read as a trend rather than a single day’s roll of the dice.
What’s actually worth tracking over time
Not every part of a visibility report needs the same testing schedule. Crawler access rules and structured data validity are close to binary and don’t drift without a deliberate change — a free scan covers that stable half in one pass, and an occasional check is usually enough. Citation rate and recommendation frequency are sampled from a live, variable process, and they only turn into a trend worth acting on after repeated runs across phrasings and models. Checking each of those on its own schedule, rather than watching one blended number, is what actually shows whether something got better, got worse, or just landed on a different roll this time. The full checklist behind both the stable and the sampled half is laid out in the methodology.
A single score is a convenient thing to put on a dashboard. It’s a poor tool for deciding what changed and why, because it can’t distinguish a real fix from a source that finally answered after failing to reach it last time, or from a model giving a different recommendation-question answer than it did a week earlier. The parts underneath the number are messier than the number itself — and they’re the ones that actually explain a change.
Frequently asked questions
Why do two scans of the same unchanged site sometimes score differently?
Part of a visibility score comes from stable, binary facts — can an AI crawler reach the site, does structured data validate. Another part comes from asking a live model a question and reading its answer, and models don't return the same answer twice on an open-ended prompt. The stable part shouldn't move between scans. The sampled part can, and a report that blends both into one number can shift even when nothing on the site changed.
What's the difference between a metric that reads 'not found' and one that reads 'not measured'?
'Not found' means the check got an authoritative answer and the site genuinely has none of whatever it was looking for — no structured data, no entry in a given index. 'Not measured' means the check never got an answer at all: a request timed out, a service it depends on returned an error, or the run ran out of time before reaching it. Collapsing both into the same blank or zero tells a reader 'confirmed clean' when the honest answer is 'nobody looked.'
Is one AI visibility score enough to track month to month?
Not on its own. A single blended number can move for reasons that have nothing to do with the site — a source that failed to answer this time but not last time, or a model giving a different response to the same recommendation question. Tracking the components that make up the score separately shows which one actually moved and why, which a single figure cannot.
How many times does a citation or recommendation check need to run before it means anything?
More than once. A model asked the same open-ended question minutes apart can return two different answers, so a single response is a sample of one. A reliable read needs the same questions asked repeatedly, phrased a few different ways, across more than one model — not a single favorable or unfavorable answer treated as the final word.
Which AI-visibility metrics are stable enough to check occasionally, and which need repeat testing?
Crawler access rules and structured data validity are close to binary and don't drift on their own — checking them occasionally is usually enough. Whether a model cites or recommends a business is sampled from a live, variable process, so it only becomes a trend worth acting on after repeated testing across phrasings and models, not a single run.
