Our scan reads one page of your site. Here is why, and what that costs.
A Biencité AI visibility scan reads your homepage and a handful of public files sitting at the root of your domain: robots.txt, the sitemap, llms.txt, and the files meant for AI agents. Nothing else. That perimeter is deliberately narrow, it takes about thirty seconds, it requires no account, and it changes nothing on your site. It therefore sees part of your site, and this article says exactly which part.
The choice rests on a property of the web rather than on technical convenience. The Robots Exclusion Protocol, standardised by the IETF as RFC 9309 in September 2022, requires that crawler rules be reachable in a file named /robots.txt at the top-level path of the service. A root file never describes a single page. It commits the whole host at once.
Why read the root of a domain instead of every page?
The root of a domain holds the decisions that apply to every page simultaneously. A robots.txt that turns GPTBot away turns it away from four thousand catalogue pages exactly as it does from your homepage. A missing llms.txt is missing for the entire domain. RFC 9309 adds that crawlers should not use a cached copy of robots.txt for more than 24 hours unless the file is unreachable, so what your root declares today is re-read quickly, and it applies everywhere.
The llms.txt file follows the same logic of scope. Jeremy Howard proposed the format in September 2024, and the specification published at llmstxt.org, updated on August 10, 2026, places the file at the root of a site or at the root of a path, covering the pages beneath it. Reading the root of a domain means reading its foundations.
What does the scanner actually read, and how long does it take?
The Biencité scanner fetches your homepage exactly as your server returns it, then a short list of public files on the domain, and the report lands in about thirty seconds. No account is required for a free scan. Here is what falls inside the perimeter.
- The HTML of your homepage as a crawler receives it, before any JavaScript runs in a browser.
- robots.txt, which states which AI crawlers you admit and which you turn away.
- The sitemap, which tells engines what addresses exist on your domain.
- llms.txt, the reading guide written for language models.
- The agent files, which describe what an AI agent can understand and trigger on your site.
Why is read-only a constraint rather than a sales precaution?
Read-only means nothing is installed, nothing is written and no access is requested from you, and that rule costs information. The scanner holds no account on your site, so it sees none of your analytics, none of your admin, and nothing behind a password. An audit that demands administrator access would see more than this one does.
What you get in return is verifiability. Everything the report shows, you can check yourself from a browser by opening the same addresses. A measurement backed by no private data is a measurement you are able to contradict, and that is the only serious reason to accept seeing less.
What does this perimeter fail to see?
This perimeter does not see your deep pages, and deep pages are exactly where AI engines go for citations. Similarweb published "The 2026 Generative AI Landscape" using data from June 2025 to May 2026: 65% of the URLs cited by ChatGPT sit two or three folders deep in a site, with depth two alone accounting for 41.7% of citations. The same report finds that 58.8% of the referral traffic ChatGPT sends lands on a homepage instead.
That gap draws the limit of the diagnosis precisely. A scan of the homepage and the root measures the entry conditions of a domain, not the citability of each product page or each article. Root signals hold for the whole site; the markup and structure of one catalogue page hold only for that page. A site whose homepage passes every signal can still publish eight hundred product pages carrying no structured data at all.
The honest answer is not to promise an exhaustive crawl. It is to state what the measurement covers, then treat the rest as continuous editorial work, page by page, running on a different clock. The article "How long before you see an effect on your AI visibility?" sets out those two clocks, the one for a technical fix and the one for regular citations.
What happens when a file does not answer?
A signal the scanner could not verify is reported as not measured and drops out of the calculation instead of scoring zero. Every signal therefore has three possible states, passed, missing or not measured, and the report states how many signals were actually verified out of the 19 in the framework. An unreachable file, a timeout, a momentary server error: a failure to measure is not an observed absence.
That distinction changes how a report should be read. A score obtained on sixteen verified signals does not compare with the same score obtained on nineteen: the first is a partial snapshot, worth repeating once your server recovers. Treating the two as equivalent produces a false score, always in the direction that flatters whoever is selling the audit.
Why does a signal you cannot reach count against nobody?
Shopify, Wix, Webflow and Framer lock certain files at the root of the domain, and no customer setting unlocks them. Those signals are marked out of reach on your platform rather than counted as gaps, and they never appear in your action plan. The principle fits in one line: you will never be handed advice your site cannot act on. The article "Shopify, Wix, Webflow, Framer: what your platform actually lets you do for AI" covers what each one still leaves open.
When will this diagnosis teach you nothing?
This diagnosis teaches you nothing until your problem is genuinely a problem of foundations. Four situations come up often, and in each of them your budget belongs somewhere else.
- Your site is three pages of introduction. The foundations will be in order quickly, and nothing will be left to cite. Write first, measure afterwards.
- You already know your content only appears once JavaScript has run. The verdict is decided in advance, and the useful spend sits with your developer.
- Your difficulty is editorial rather than technical: your pages are readable, structured, and nobody reads them. A root-level scan says nothing about that.
- You want to know what AI engines say about your company today. That is a different measurement, weekly and repeated, described in the article "Asking ChatGPT whether it knows you measures nothing".
The full list of the 19 signals and the way the score reads are published on the Biencité method page, at https://www.biencite.fr/en/methode. The weight of each signal is not published: those weights move with the practices of generative engines, and publishing them would invite people to optimise a score rather than real visibility.
Frequently asked questions
Does the scan change my site?
No. The scanner reads your homepage and a few public files on the domain, exactly as an AI crawler would. Nothing is installed, nothing is written, and no account is needed to run a free scan.
Why not scan every page of my site?
Because the decisions being measured are taken at the root of the domain and hold for every page at once. A root-level scan measures the foundations. Quality page by page is continuous editorial work, running on a different clock.
Is a good homepage score enough to get cited?
No. The score measures technical eligibility, never a citation. Similarweb finds that 65% of the URLs cited by ChatGPT sit two or three folders deep, so your deep pages have to earn the citation too.
What does a not measured signal mean in my report?
That the scanner could not verify it, because of an unreachable file, a timeout or a server error. The signal leaves the calculation instead of scoring zero, and the report states how many signals were genuinely verified.
- ai visibility scan
- read only audit
- robots txt
- llms txt
- geo