Nedim Mehić
Back to blog

Measurement · · 5 min read

AI Citations Are the New Vanity Metric

A citation, a recommendation, a visit, and a customer are four different outcomes. Microsoft says as much in its own documentation. Report them separately.

Every AI visibility tool now opens on a large number. You were cited 4,000 times this month.

Good. Now what?

I ask that question in a lot of reporting calls and the honest answer is usually silence, followed by a slide about brand awareness. That is the tell. A number nobody can act on is not a metric, it is decoration, and this particular decoration is about to become the most over-reported figure in the industry.

Four different outcomes wearing the same name

A citation, a recommendation, a visit, and a customer sit at four different points on a chain, and each step loses most of the volume above it.

Cited means your URL was among the sources an answer drew on. It may appear in a collapsed footnote nobody expands. Pew's browsing study found that links inside AI summaries were clicked in roughly 1% of visits. Being in the source list is very far from being chosen.

Recommended means the answer named you as the pick. This is a categorically different outcome and it is the one most people think they are measuring when they look at a citation count. You can be cited heavily in an answer that recommends a competitor, and cited as the example of what not to do.

Visited means a person actually arrived. Partly measurable through referrers, badly, since plenty of assistant traffic arrives without one.

Customer means revenue. This is the only number in the list that pays anybody.

Collapsing these four into one headline is how a reporting deck ends up looking excellent while a pipeline does not move.

Microsoft says this out loud

This is not my framework. It is close to what the platform documentation already states.

Bing Webmaster Tools' AI Performance reporting is explicit that its citation figures reflect overall citation patterns and do not indicate ranking, authority, or the role of a given page within an individual answer. The newer Citation Share metric is described as observational rather than a ranking system or competitive scoreboard.

The vendor with the most access to this data is telling you, in its own help documentation, not to read the number as a measure of standing. That is worth more than most third-party methodology claims. The details are in Microsoft's AI Performance documentation and the public preview announcement.

Why citation counts inflate so easily

Three reasons, all of them mundane.

Query mix. Citation volume follows the questions people ask, and most questions in your category have no commercial intent. A thousand citations on definitional queries and none on the comparison query that decides a purchase is a worse month than the reverse, and the headline number will report it as ten times better.

Sampling and modelling. Nobody observes every answer. Third-party tools run a panel of prompts and extrapolate. Change the prompt set and the number moves, with no change in the world. Always ask what the prompt list is and who chose it.

Answer position. Being the source of the first sentence and being source eleven in a list are the same event in most counts. They are not the same event for a reader.

How to report it honestly

Break the single number into the chain and report each step with its own source.

  • Citation presence. Platform reporting where it exists, Bing Webmaster Tools first. Track it by topic, not in aggregate, so you can see which parts of your category you are absent from.
  • Recommendation share. Track a fixed list of the prompts that actually precede a purchase in your market. Twenty is plenty. Record whether you are named, and whether you are named first. Same prompts every month, or the trend line is meaningless.
  • Arrival. Referrer data from the assistants that send it, plus direct and branded search as a proxy for the ones that do not. Expect this to be lossy and say so in the report.
  • Revenue. Self-reported attribution on your forms. One open question about how someone heard about you will outperform every dashboard you can buy.

Four rows, four sources, four different confidence levels stated plainly. It is a less impressive slide and a far more useful one.

The filter

Before any number goes on a report, I ask one question.

What would I do differently if this doubled?

If citation volume doubled, would you change the roadmap? Probably not, unless you knew which topics drove it. If recommendation share on your twenty buying-intent prompts doubled, would you change anything? Yes, immediately, and you would know exactly which pages earned it.

That difference is the entire argument. One number is a scoreboard for a game nobody is paying you to win. The other tells you where to spend next quarter.

Being visible in AI answers is genuinely worth pursuing. Just celebrate it in proportion to the business value you can actually demonstrate, and be the person in the room willing to say which part of the number is still a guess.

Keep reading

Work with me