Analytics / AI Search

AI Visibility Analytics: What You Can Actually Measure

10 min read By the Neurobird team
A printed line chart on matte paper with one line plotted and a second deliberately left blank, representing what can and cannot be measured
Short answer

Four things about AI visibility are genuinely measurable: whether AI crawlers can reach you, whether you appear in generated answers for a fixed question set, where you rank across discovery registries and the open web, and real adoption from registries that publish install and dependency counts. One thing is not measurable by anyone: how often agents searched for your category. Position is the most valuable of the four because it cannot be reconstructed later, so recording it today is the only part you cannot buy back.

Most AI visibility dashboards lead with the one number nobody can produce. Knowing which metrics are verifiable, and which are a language model's guess in a confident font, is most of the work of building reporting that survives a sceptical question from your CFO.

Start with the number nobody can give you

The first question every brand asks about AI visibility is how many people, or agents, are asking about their category. It is the natural question, it maps onto the search volume everyone spent fifteen years optimising against, and it cannot be answered.

Answer engines do not publish query logs. Discovery registries do not publish query logs. There is no equivalent of a keyword planner because there is no party with both the data and a reason to share it. So when a dashboard shows you a monthly figure for how often AI systems were asked about your category, that number came from one of two places: the vendor's own traffic, or a language model asked to estimate. Neither is a measurement of your market.

If a metric cannot be reproduced by you, from a source you can check, it is a projection with a confident font.

This matters more than it sounds, because the unmeasurable number is also the flattering one. Nobody is motivated to check a figure that makes their category look large.

The four things that are genuinely measurable

All four can be verified independently, which is the test worth applying to any metric somebody wants to charge you for.

1. Crawler access, from your own logs

The cheapest and most neglected. Your server already records every request, including the ones from GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and Claude-SearchBot. If those user agents never appear, nothing downstream can possibly work, and no amount of content strategy fixes it.

This is a five-minute check with a real answer at the end. The most common cause of zero AI visibility is still a robots.txt written in 2023 to keep scrapers out, doing exactly what it was told.

2. Citation presence, by asking

You cannot see how often a question was asked, but you can absolutely see what the answer is. Write down the twenty or thirty questions a buyer would actually ask, ask them on a schedule across the engines that matter to you, and record whether your brand appears and who else does.

The discipline that makes this a metric rather than an anecdote is fixing the question list before you look at results, and never editing it because a question is unflattering. A list curated after the fact measures your curation.

3. Position, across registries and the open web

This is the closest thing to a classic rank tracker and it is the most underused. For a fixed set of phrasings, record where you appear in each public discovery registry and in web results, every day.

Two properties make position the most valuable metric in this list. It is reproducible, so anyone can re-run the query and get the same answer. And it cannot be back-filled: if nobody recorded where you ranked last month, that month is gone permanently. It is the one measurement where starting late has a permanent cost, which is the real argument for beginning before you have a strategy.

4. Adoption, from sources that publish real counts

The one place hard numbers genuinely exist. Package registries publish download counts. Code hosts publish stars, forks and dependants. Container registries publish pull counts. These are counts other people maintain, free to read, and checkable with a single request.

They are not a proxy for demand and should not be presented as one. A download is not a user: registry figures include continuous integration, mirrors and bots, sometimes overwhelmingly. Reported honestly, as installs rather than customers, they are the most solid signal available.

Building a metric you can defend

Four rules, each of which comes from a way these systems go wrong in practice.

RuleWhy
Separate an outage from an absenceIf a source fails to answer, record nothing. Storing a failed request as a zero tells you that you lost a ranking on a day when the only thing that happened was a timeout.
Fix the question set in advanceQuestions chosen after seeing results measure the person choosing.
Never blend components into one scoreThe weights in a composite are an opinion. Keep stars, installs, citations and positions separate and let a reader combine them if they want to.
Say the denominator out loud"You rank eleventh" means nothing without "of the 341 publishers we can resolve numbers for". A rate without a denominator is decoration.

What good looks like in practice

A weekly view that answers four questions, each with a source attached:

Notice what is absent: no volume, no opportunity score, no visibility index out of a hundred. Every line is a count or a position, and every line names where it came from.

The two mistakes that make an analytics programme useless

Measuring once. Answer engines are non-deterministic and registries change daily. A single run is a sample. The value is entirely in the series, which is why the setup cost is worth paying before you think you need it.

Optimising the metric instead of the thing. The best-documented finding on writing for AI answers, from the Princeton study presented at KDD 2024, is that citing sources, adding direct quotations and including named statistics all measurably improve the odds of being cited, while keyword stuffing performs worse than the baseline it was supposed to beat. The techniques that game a metric and the techniques that earn a citation point in opposite directions, which is unusually convenient.

One related caution on structured data. FAQ rich results were retired from Google search on 7 May 2026, so FAQ schema no longer buys a visual result. It remains worth adding, because clean question and answer pairs are exactly the shape an answer engine lifts, but if your reporting still counts rich results as a KPI, that KPI describes a feature that no longer exists.

Where to start this week

  1. Grep your access log for the AI crawler user agents. If they are absent, stop and fix robots.txt first.
  2. Write the twenty questions. Do it before you look at anything, and keep the file.
  3. Record position today, even manually. Today's snapshot is the only part of this you cannot obtain later.
  4. Pull your real adoption counts from the registries that publish them, and write them down as installs, not users.

None of that requires a platform, and doing it by hand for a month teaches you which parts are worth automating for your specific category.

Can you measure how often AI agents search for your product?

No, and any vendor offering that number cannot either. Discovery services and answer engines do not publish query logs, so an agent search volume figure is either the vendor's own test traffic or a model estimate. Position and adoption are measurable; demand is not.

What can you actually measure about AI visibility?

Four things, all of them checkable. Whether AI crawlers can reach you, from your own server logs. Whether you are cited in generated answers, by asking the questions and recording the result. Where you rank across discovery registries and the open web for the phrasings people use. And real adoption, from the package registries and code hosts that publish install and dependency counts.

Is citation share a real metric?

It is real if you define the question set first and keep it stable. Citation share is the proportion of a fixed list of buyer questions where your brand appears in the answer, measured on a schedule. It becomes meaningless the moment the question list is chosen after seeing the results.

Why do AI visibility tools show different numbers for the same brand?

Because they ask different questions, at different times, from different locations, and answer engines are non-deterministic. A single run is a sample, not a measurement. Only a fixed question set sampled repeatedly over time produces a trend you can act on.

How often should AI visibility be measured?

Daily for position, because history cannot be reconstructed after the fact. Nobody can sell you last month's rankings if nobody recorded them, which is the one genuine reason to start measuring before you need the answer.

Want to see where you stand right now?

The free audit checks crawler access, structured data and answer-shaped content in seconds, and tells you what an AI system can currently learn about your brand.

Run the free audit

Related Neurobird products

Neurobird also builds focused operational software for specialised industries. The full catalogue is at neurobird.com/solutions.