SEO · GEO · AEO

AI visibility audits: how to measure mentions, citations and gaps

A reproducible method for evaluating how a brand appears in AI answers without turning volatile results into an absolute score.

Executive brief

Key takeaways

  • Choose questions tied to real journeys and decisions.
  • Record the full context of every collection.
  • Separate mention, recommendation, citation and factual accuracy.
  • Use trends across repeated runs, not one answer.

An AI visibility audit measures how a brand, product or content appears in generative answers to important questions. It does not try to find a permanent position because these answers vary. The goal is to build a reproducible baseline, locate gaps and guide experiments.

The unit of analysis is no longer only “keyword and position.” It includes question, context, model, answer, mentioned entity, cited source, accuracy and the next action suggested to the user.

Start with the journey, not the tool

Hundreds of prompts unrelated to business decisions create volume without direction. Organize questions by stage.

Problem discovery

These questions help users name a need: why a website is slow on mobile, how to know whether a page is indexable, or how to evaluate presence in AI answers.

Solution exploration

The user compares methods or categories: website audit tools, the difference between technical audit and conversion analysis, or whether a B2B company needs SEO, GEO or AEO.

Alternative comparison

Brands, products and criteria become explicit: which tools combine SEO, performance and accessibility; how to choose a website diagnostics platform; alternatives to generic audit reports.

Decision and objection

Questions address trust, integration, risk and process: can a public audit prove lost revenue, does the report include evidence and an action plan, and who should implement recommendations?

This map prevents easy but low-value informational questions from dominating the audit.

Define a reproducible protocol

For every collection, record:

  • exact question text;
  • engine, product and model when available;
  • date and time;
  • language and country;
  • new session or contextual conversation;
  • enabled resources, such as web search;
  • preserved answer or evidence;
  • links and sources shown;
  • reviewer and classification rule.

Without these fields, two answers cannot be compared confidently. A change may come from content, a system update, location, conversation history or normal generation variance.

Classify results along separate dimensions

Presence

Does the brand appear? Record absence, incidental mention, list inclusion, explicit recommendation or primary emphasis. Do not reduce all outcomes to one “yes.”

Attribution

Does the answer connect a claim to the brand? Is there a clickable URL? Does the link lead to the correct page? Does the cited source support the statement?

A brand may be mentioned without a source. A page may be cited without prominent brand recognition. These are different outcomes.

Accuracy

Are the name, product, capabilities, audience and limitations correct? Incorrect information may require a clearer page, consistent entity data or updates to third-party profiles. It may also be a model error with no direct corrective action.

Intent coverage

Does the brand appear at the right journey stage? Being cited in a broad definition differs from appearing in a tool comparison or buying recommendation.

Source quality

Record which domains support the answer. Official documentation, media, directories, communities and affiliates play different roles. The goal is to understand where the system forms its view.

Use a repeated sample

One run is one observation. Repeat the question set in controlled sessions to identify patterns. Frequency depends on program size and topic volatility, but the protocol should stay comparable.

QuestionRunsPresenceCitationAccurateObservation
Website audit tools5212Brand appears only in answers with web search
How to prioritize findings500Brand content does not cover the intent
Audits and revenue5323Product limitations described correctly

These values would be sample counts, not market percentages.

Connect observations to the site

When a gap appears, investigate verifiable causes:

  • whether a page answers the question;
  • whether the answer is available in HTML;
  • whether title and introduction make the thesis clear;
  • whether authorship, date and entity data are consistent;
  • whether evidence stays close to claims;
  • whether relevant pages link to the content;
  • whether the analyzed system can access it;
  • whether external sources describe the brand consistently.

Do not assume a page change will produce a citation. Record a hypothesis and test it.

Prioritize commercially meaningful gaps

An absence matters most when the question represents a real audience journey, the brand is credible on the subject, a plausible technical or editorial action exists and the result can be observed again.

Broad questions with apparent scale may be less useful than specific comparisons close to a decision. Prioritization should consider relevance, not only mention frequency.

Metrics that can help

  • Sample presence rate: percentage of controlled runs where the brand appears. Do not call it absolute model share.
  • Citation rate: percentage with a source or link attributable to the brand. Separate homepage links from the page supporting the answer.
  • Accuracy: percentage of present answers without a material entity or offer error. Define “material” before reviewing.
  • Journey coverage: number of stages and intents with enough presence to analyze.
  • Traffic and assisted conversion: sessions, engagement and conversions when identifiable references exist. Not every mention creates a link, and not every link preserves attribution.

In June 2026, Google announced tests of dedicated generative feature reports in Search Console, initially for a subset of sites. Confirm availability and fields in the account before making it part of a standard process.

State the limitations

Answers vary and can be personalized. Models and indexes change. Some interfaces do not show every source. Location and language affect output. Large-scale automation may conflict with service terms. A prompt sample cannot represent every possible question. Absence does not prove lack of influence, and presence does not prove preference, trust or conversion.

Turn the report into an experiment

Observation: Remountly did not appear in five runs about audit prioritization. Site evidence: there is no pillar guide dedicated to turning findings into a backlog. Hypothesis: publishing an original method and connecting it to technical content improves topic coverage. Action: publish the guide, create internal links and repeat the sample after indexation.

This format is more useful than “improve GEO.” It explains the gap, why the action makes sense and what remains uncertain.

AI visibility is a program of observation and learning. With a stable method, volatility becomes an explicit part of the evidence.

Direct answers

Frequently asked questions

Is there an official GEO score?

There is no universal score accepted across generative engines. Tools may create proprietary indexes, but their method and limitations should be visible.

How many questions should I test?

Start with a small set covering discovery, comparison, objections and decisions. Expand only when the protocol is stable and the team can review answer quality.

Does an unlinked mention have value?

It may contribute to recognition or consideration, but it is different from a verifiable citation or traffic. Classify each outcome separately.