What an AI visibility scan actually measures, and what it cannot

An AI visibility scan measures technical eligibility: whether a site can be read, then cited, then used by a generative AI assistant. At Biencité, that measurement takes the form of 19 signals across 3 axes, checked read-only against the home page and the domain's public files. An AI visibility scan does not measure your reputation, does not report what ChatGPT already says about your company, and promises no citation.

The commercial stake behind that measurement is documented. Ahrefs compared 300,000 keywords between December 2023 and December 2025 and published the result on February 4, 2026: when an AI Overview appears at the top of the results page, the top organic position receives 58% fewer clicks on average. The traffic does not vanish, it moves into the generated answer. Appearing inside that answer becomes a revenue question.

What does an AI crawler actually read when it lands on your site?

A generative AI crawler reads the HTML your server returns, and nothing else. Vercel analysed roughly 1.3 billion AI crawler requests across its network and published the findings on December 17, 2024: none of the major answer crawlers, GPTBot and OAI-SearchBot at OpenAI, ClaudeBot at Anthropic, PerplexityBot at Perplexity, execute the JavaScript on the pages they visit.

That limit explains why one signal on the Visible axis carries so much weight. A site whose content only appears once JavaScript runs in the browser is an empty shell to these crawlers: they receive a page with no text, read it, and move on. Google is the exception, because its rendering engine executes JavaScript before indexing, and Gemini inherits that capability by running on Google's infrastructure. The result catches a lot of owners off guard: a page can rank well on Google and be completely absent from ChatGPT's answers, which is also why an audit priced purely on Google metrics misses the point, as the article « How much does an AI visibility audit cost in 2026? » sets out.

Why is a single overall score not enough to describe a site?

A single overall score flattens three distinct problems into one number, when those three problems call for three different responses. Biencité therefore splits the 19 signals across 3 axes that are read separately.

  • Visible, 5 signals: secure connection, AI crawlers allowed in, content readable directly, essential page information, usable site map. This is the entry condition, and everything else depends on it.
  • Citable, 8 signals: structured data, content structured for citation, sharing preview, dedicated reading guide for AI, declared AI usage preferences, structured key facts, content ready for voice, linked entities and profiles. This is what turns a readable site into a reusable source.
  • Agent-ready, 6 signals: agent-readable version, capability card, skills index, authentication framework, exposed service catalog, agent access guide. This is the layer almost no site has yet.

A site can be perfectly visible and still be uncitable. It happens constantly: the content is in HTML, the crawler reads it without trouble, but nothing on the page says who is speaking, about what, or on what authority. The AI sees the site without being able to recommend it. The action plan that follows a scan then covers the Citable axis only, and leaving the Visible axis alone saves real time.

Why does the verdict not follow directly from the score?

The verdict follows a cascade, not an average. A site that AI assistants cannot read is classed « invisible to AI » whatever the rest of its score, even if it ticks advanced signals on the other two axes. The four verdicts run in order: invisible to AI, visible but not citable, readable and citable, ready for AI agents.

The cascade exists to prevent a lie by arithmetic. A site closed to AI crawlers but carrying flawless structured data would post a flattering number under an averaging model, while appearing in no generated answer at all. The verdict states the operational truth instead: while the door is shut, the interior decoration does not count.

What happens when a signal cannot be measured?

An unmeasured signal drops out of the calculation instead of scoring zero. An unreachable file, a timeout, a momentary server error: the scanner marks the signal as unmeasured, and the report states how many signals were genuinely checked out of 19. A failure to measure is not an observed absence, and conflating the two produces a wrong score that happens to favour whoever is selling.

The practical consequence matters when you read a report. A score of 62 based on 17 verified signals is not comparable to a score of 62 based on 19. The second is a complete photograph, the first is a partial one worth retaking once the server recovers.

Why does your platform change your result?

Shopify, Wix, Webflow and Framer lock certain files at the domain root, and no customer-side setting unlocks them. Those signals are marked impossible on your platform rather than counted against your site, and they never appear in your action plan. The principle is straightforward: you never receive advice your site cannot act on.

That choice carries a deliberate commercial cost. A longer list of recommendations impresses a prospect more, and plenty of audits are happy to produce one. An unactionable recommendation is not a gap to close though, it is a platform constraint, and presenting it as a defect wastes everyone's time.

Should you let AI crawlers in, or block them?

OpenAI publicly documents separate crawlers for separate purposes, which makes the decision finer than a simple open or shut. OAI-SearchBot feeds ChatGPT's search answers, and OpenAI states that a site disallowing it will not appear in those answers. GPTBot collects content for training foundation models. Blocking both in one gesture means withdrawing from the answers in order to keep your text out of training.

biencite.fr runs the position it recommends: answer crawlers allowed, pure training crawlers refused, usage preferences declared in a readable form. That file carries no enforcement power, which is the subject of the article « Your robots.txt is not a lock: AI bots walk past it, and OpenAI says so ». A scan checks what your site declares, not what the crawlers then do with it.

When is an AI visibility scan not worth running?

An AI visibility scan is worthless while there is nothing to cite. Four situations come up regularly, and in each of them the money belongs somewhere else.

  • A three-page site with no substantive content. Making a brochure citable produces nothing: an AI cites an answer, not an address. Write first, measure second.
  • A business with no trace anywhere online. Generative AI assistants cross-reference sources, and a perfectly structured site that nobody mentions elsewhere is still an isolated site.
  • A budget still owed to more urgent foundations. A site that is down half the time, or an empty business listing, gets fixed before the AI layer.
  • A heavily locked platform on a project migrating within six months. Waiting for the migration beats paying twice for the same work.

What will you still not know after a scan?

A scan does not tell you what AI assistants currently say about your company. That is a separate measurement, citation tracking, which queries ChatGPT, Claude, Perplexity, Gemini and Le Chat every week to see who gets cited, you or your competitors. A scan also does not judge the editorial quality of your writing, does not measure your backlinks or your reputation, and predicts no citation.

The weight of each signal is not published, and there is a reason for that. Those weights shift as generative engines change their practices, and publishing them would invite optimising the score rather than the underlying visibility. An AI visibility score is compared against itself over time, never against another site's. The full list of signals is published on the Biencité methodology page, because a measurement nobody can inspect does not deserve trust.

Frequently asked questions

Does the scan modify my site?

No. The scanner reads your home page and a handful of public files on the domain, exactly as an AI crawler would. Nothing is installed, nothing is written, and no account is needed to run a free scan.

Does a high score guarantee ChatGPT will cite my company?

No. A high score means your site meets the technical conditions without which a citation is highly unlikely. What AI assistants then cite depends on your reputation, your sources and the question asked. No serious tool can promise a citation.

Why 19 signals and not fifty?

Because each retained signal maps to a capability that a generative AI actually uses to read, cite or act on a site. Padding the list with criteria that change nothing would inflate the report without improving visibility.

Why are the signal weights not published?

Because they shift as generative engines change their practices, and publishing them would invite optimising the score rather than the underlying visibility. The signals themselves are all published on the methodology page.

My site runs on Shopify, can it still improve?

Yes, on most of the signals. Shopify locks certain files at the domain root: those signals are marked impossible on your platform, drop out of your action plan, and are not counted as failures.

Topics
  • ai visibility scan
  • geo
  • ai audit
  • methodology
  • citation eligibility
What an AI visibility scan actually measures, and what it cannot