5,179 queries to five AI systems, over four months, for eleven websites.
The result was 4.1 per cent. The correct figure is 6.3 per cent.
The difference had nothing to do with the AI systems. It was our own measurement.
Short answer: When measuring AI citations, the query is rarely what goes wrong. What goes wrong is the frame around it — the denominator, the choice of systems, the detection. This piece shows both on our own data: what we measured, and where we fooled ourselves doing it.
The thesis: AI answers are random, measuring is pointless
You hear the argument often and it sounds compelling. A language model rolls the dice on every answer. Ask the same question twice and you get different wording and different sources. Measuring that means measuring noise.
On top of that, the systems change weekly. A new model, a new search integration, a different approach to attribution — and the time series is void without anyone noticing.
If that holds, every citation measurement is busywork.
The antithesis: our data says otherwise
We checked. For one project we hold 215 combinations of phrase and system that were each measured at least five times, many of them thirteen to eighteen times.
For 199 of those 215 pairs the result was identical across every single measurement. Cited stayed cited, not cited stayed not cited. Only 16 pairs — 7.4 per cent — ever switched between measurements at all.
That is the opposite of random. The wording of the answer does fluctuate. Whether your domain appears as a source, as a rule, does not.
And the systems differ clearly and consistently from one another:
| System | Pairs checked | Citations | Rate |
|---|---|---|---|
| Gemini | 1,649 | 125 | 7.6% |
| ChatGPT | 588 | 39 | 6.6% |
| Perplexity | 588 | 36 | 6.1% |
| Claude | 556 | 12 | 2.2% |
Claude cites most sparingly, Gemini most generously — over four months, across eleven websites. A factor of three cannot be explained by noise.
What actually goes wrong: the frame, not the query
Here comes the uncomfortable part. Our own overall rate read 4.1 per cent — 212 citations out of 5,179 checks. That number was wrong, for a reason that can lurk in any citation measurement.
Error 1: a denominator that cannot produce hits
Of those 5,179 checks, 2,007 went to Google AI Overviews. Not one of them ever registered a citation. Not rarely — never, across two months.
As long as those 2,007 sit in the denominator, they depress the rate while being structurally unable to add to the numerator. Without them, the same measurement reads 6.3 per cent — 212 out of 3,381.
A third of a difference, purely from the composition of the denominator. Nothing about the websites had changed.
Error 2: detection that finds the wrong thing
The second error is trickier. For those same Google AI Overviews our measurement counted 282 mentions — the brand appeared, but without a link.
Opening the raw data showed this:
was_cited: 0
cited_url: NULL
raw_answer: Reisebericht Tobago — https://www.my-travelworld.de/tobago/...
The domain is in the text. It still does not count as cited. The reason: our parsing code put the AI Overview text and the organic search results into the same container, and the mention detection then searched all of it.
So the domain was found because it ranks organically — not because an AI mentioned it. Those 282 mentions measure classic SEO to an unknown degree and sell it as AI visibility.
Error 3: two names for the same system
Along the way we found two identifiers for Google AI Overviews: one with 7 rows up to early June, one with 2,007 rows from mid-June. A rename without migrating the old data.
Anyone evaluating a time series across that point while filtering on only one of the two gets a break that looks like a collapse.
The synthesis: four facts without which a number says nothing
Both sides of the argument are wrong. Measurement is not worthless — it is more stable than its reputation suggests. But it is not self-explanatory either. A citation figure needs four facts beside it, or it is not a statement.
1. The denominator, broken down by system
Not “5,179 checks”, but how many went to which system. That is the only way a system that systematically returns nothing becomes visible.
2. How the phrase list is composed
Does it rotate automatically because it follows your Search Console data? Then the yardstick shifts every month, and two figures from two months are not comparable even though they look alike.
3. What counts as a citation
A linked source under the answer? The domain anywhere in the text? The brand name without a link? Those are three different things with three different values. Our error above happened at exactly that boundary.
4. The period and the measuring frequency
Weekly measurement of twenty phrases across five systems gives 400 checks a month. Daily gives 3,000. Same website, seven times the denominator.
What we did about it
Detection for Google AI Overviews is being reworked and the duplicate identifier merged. Until then, for our own figures: the four systems with reliable detection are Gemini, ChatGPT, Perplexity and Claude. The values in the table above come from those four alone.
How we got there is the actual point of this article: the error only surfaced because someone opened the raw data instead of believing the metric. An interface that only shows “4.1%” would never have made it visible.
Conclusion
AI citations can be measured, and the result is remarkably stable — 92.6 per cent of our repeatedly checked pairs always returned the same outcome. Anyone claiming it is all random has probably not measured repeatedly.
What is unreliable is not the AI’s answer. It is the number made from it when nobody names the denominator.
So ask of every tool — including ours: out of how many checks, on which systems, with what definition of “cited”? A vendor without an answer is selling a percentage with no meaning.
Frequently asked
How stable are AI citations really?
More stable than expected. Across 215 phrase-and-system combinations each checked at least five times, the result stayed identical for 199 of them. Only 7.4 per cent ever switched. What fluctuates is the wording of the answer — not whether a domain shows up as a source.
Which AI system cites most often?
In our data Gemini at 7.6 per cent, then ChatGPT at 6.6 and Perplexity at 6.1. Claude sits well below at 2.2 per cent. The figures come from 3,381 checks across eleven websites between April and August 2026 and cover German-language queries — a snapshot, not a universal benchmark.
Why is the citation rate alone not a usable metric?
Because its denominator moves. In our own case the rate read 4.1 per cent while 2,007 checks sat in the denominator that were structurally unable to produce a citation. Without them it was 6.3 per cent. A third of a difference, with nothing having changed on a single website.
What separates a citation from a mention?
A citation is a linked attribution: the AI names your page as a source and points to it. A mention only names the brand in the text, without a link. Both have value, but different value — and folding them into one number makes it useless. More in the glossary under AI Citation Tracking.
How often should you measure?
Less often than most people think. If 92.6 per cent of pairs return the same result for weeks, a daily check adds almost no insight — it only multiplies the cost, because a paid model call sits behind every check. Weekly is enough for monitoring; before and after a change is the moment when an extra measurement actually answers something.
Quick quiz
Four questions on what a citation figure says — and what it does not.
1. The same phrase is checked 15 times on the same system. How often does the result change?
b is correct. Across 215 repeatedly checked pairs the result stayed identical in 199 cases. The wording of the answer fluctuates; the choice of sources barely does.
2. A rate drops from 6.3% to 4.1%. What is the most likely cause?
c is correct. Exactly what happened to us: 2,007 checks on a system that never returned a citation pushed the overall rate down by a third — with no change to any website.
3. The domain is in the answer text, but the tool reports “not cited”. What do you check?
a is correct. Our own error: the parsing code put the AI answer and the organic results into the same text block. The domain was found because it ranks, not because an AI named it.
4. What separates a citation from a mention?
b is correct. Both have value, but different value. Folding them into one number makes it useless — and that boundary is exactly where our measurement error occurred.