Your site is invisible to AI: the five most common causes
Five causes explain almost every case of a site that never appears in AI answers: its content is not present in the HTML served to bots, its infrastructure blocks answer engines without anyone having decided to, the site never declares which company it describes, it has no substantive content worth quoting, and its answer exists but sits buried in a page that isolates no single passage. They are presented here in order of the cost of leaving them alone, not in a measured order of frequency: we do not publish a ranking we have not measured.
The stakes are easy to understate because nothing in your analytics announces the loss. Ahrefs measured across 300,000 keywords that when an AI overview appears, the click-through rate of the number one organic position is 58% lower (study published February 4, 2026). A site absent from the answer loses visibility its Google ranking no longer makes up for.
Can AI actually read the content on your site?
The first cause of invisibility is content that is not in the HTML the bot receives, but assembled afterwards by the browser. An analysis Vercel published on December 17, 2024, covering more than 500 million bot requests including 569 million GPTBot requests in a single month and 370 million for Claude, concluded that none of the major AI crawlers executed JavaScript: they downloaded the files without ever running them. A client-rendered site appeared empty to them.
This point deserves an honest update, because it is moving. On September 25, 2026, a developer observed in their own server logs, across four Next.js sites, OAI-SearchBot jumping from 1,253 to 41,444 requests in a single day and GPTBot from 83 to 28,484, with signs of genuine rendering such as fetching images and stylesheets (hisashi.space, September 30, 2026). OpenAI documents nothing about this on its official bots page, and the author notes that the rendering appears to be switched on for some sites only.
The practical conclusion does not change. Making your visibility depend on JavaScript execution means betting on behaviour that is undocumented, unguaranteed and not uniform: the September 2026 observation concerns OpenAI's bots alone, and says nothing about Claude, Perplexity or Le Chat. Serving your content directly in the HTML remains the only option that depends on no single engine. Our article "What an AI visibility scan actually measures, and what it cannot" covers how this signal gets checked.
Who is really blocking AI bots on your site?
The second cause is bots being blocked, almost never where the owner assumes. The index HasData published on September 29, 2026, which measured 10,894 domains (the Tranco top 10,000 plus 1,148 news publisher sites) on July 10, 2026 and again on September 16, 2026, found a gap that robots.txt does not explain: from the same datacenter address, GPTBot received a valid response only 47.2% of the time on publisher sites, against 83.1% for an ordinary browser. From a residential address, ClaudeBot got through just 36.8% of the time.
The same index shows declaration and reality diverging in both directions. On one side, 39.5% of sites that disallow GPTBot in robots.txt served it a valid response anyway when it knocked. On the other, 5.5% of sites block bots without having declared it anywhere. Your robots.txt is not where the decision gets made: the firewall, the CDN or the host's anti-bot setting decides, often by default, and usually without anyone on the business side having weighed the trade-off.
- A protection setting enabled by default at the host or CDN, filing answer engines alongside unwanted scrapers.
- A security plugin installed for an unrelated reason, screening out unfamiliar user agents.
- A robots.txt copied from a template found online, disallowing bots whose role nobody checked.
- An attack-mitigation challenge that requires running a script, which an answer engine will not do.
Does your site say clearly which company it describes?
The third cause is a site that is readable but never declares which entity it describes. An AI asked to recommend a supplier needs to attach content to a named company with a location, a line of business and profiles verifiable elsewhere. Without structured data or linked entities, the text stays text: accurate, possibly useful, but not attachable to an organisation the engine could cite with any confidence. That is the difference between a site that discusses plumbing and a site an AI knows belongs to one specific plumbing company in Leeds, with consistent public reviews and profiles.
This is the quietest of the five, because it is invisible on screen. A human visitor works out which company it is from the logo and the footer. An answer engine infers nothing: it reads what is declared.
Is there anything on your site worth citing?
The fourth cause, and the most common on brochure sites, is having nothing to cite. A five-page site stating a trade, a service area and a contact form offers an AI no passage to extract. An engine does not cite a company because it exists, it cites a passage that answers the question asked. Without substantive content, no technical setting produces a citation, and this is the case where a measurement tool is useless: the problem is not measurable, it is editorial.
This is also the cause that argues against buying any tool yet. If your site holds no content a buyer could quote from memory, the budget belongs in writing a handful of solid pages rather than in a monitoring subscription. Measuring technical eligibility on a site with nothing to say is like inspecting the plumbing of a house with no water.
Is your answer extractable, or buried in the page?
The fifth cause is an answer that exists in a page that never isolates it. A generative engine extracts one to three sentences, rarely more. A sentence opening with "as we saw above" or "the latter" loses its meaning the moment it leaves the page, and becomes unusable. A price locked inside an image, an answer sitting two thirds down a page with no intervening heading, an argument diluted across six paragraphs: the content is there, it is not extractable.
This one is fixable without a rebuild: a heading phrased as a question, the answer in the sentence that follows, a paragraph that stands on its own. The same content becomes citable without a line of code changing.
Is a missing llms.txt a cause of invisibility?
No, and it is worth saying plainly. Ahrefs examined 137,210 domains in May 2026 and found that 97% of llms.txt files received no requests at all over the month (study published June 15, 2026). The file is cheap to produce and we publish one on our own site, but it does not belong among the causes of invisibility: a site whose content is missing from the HTML will stay invisible with a flawless llms.txt.
The distinction matters when a budget is at stake. Agent-facing files are a later layer, worth adding once access, readability, entity and substance hold, and pointless before that.
In what order should you tackle these five causes?
In the order in which they cancel each other out. Access and readability first, because a site bots cannot reach or cannot read gains nothing from the rest: that is the cascade we apply, where an unreadable site is reported invisible whatever the quality of its other signals, as our /en/methode page sets out. Entity identification next. Editorial substance and extractability last, since those are real work rather than a setting.
Your platform sometimes limits what you can fix yourself: Shopify, Wix, Webflow and Framer lock certain files at the root, and our article "Shopify, Wix, Webflow, Framer: what your platform actually lets you do for AI" covers what is feasible and what is not. The initial diagnostic is free and runs read-only, with nothing installed on your site.
Frequently asked questions
How do I work out which of these five causes applies to my site?
The first two, access and readability, can be checked technically and read-only from outside the site: request the page the way an answer engine would, and look at what comes back. Entity identification checks the same way. The last two, missing substance and an unextractable answer, need an editorial read of your pages, which no automated measurement replaces.
My site ranks well on Google, so why would it be invisible to AI?
Because the two do not read the same thing. Googlebot executes JavaScript and holds a history of your site, letting it understand pages an answer engine receives empty. A good Google ranking proves Google can read you, not that ChatGPT, Claude or Perplexity can.
Do I have to let AI bots in to be cited?
Separate answer engines, which fetch a page to answer a live question, from pure training crawlers. Blocking the former means refusing the citation you are chasing. The trade-off is made bot by bot, and the setting usually lives with your host or CDN rather than in your robots.txt.
How long does fixing these causes take?
Access and readability resolve in days, sometimes in a single host setting. Entity identification sits on the same scale. Editorial substance and extractability take weeks, because they mean writing or rewriting pages.
Can a site suffer from several of these at once?
Yes, and that is the usual case. A recent brochure site built on a quick site builder often combines client-side rendering, absent structured data and content too thin to quote. Working through them in order matters because fixing access is what makes the later fixes worth anything.
- invisible to ai
- ai visibility
- geo
- why chatgpt does not cite my site
- ai crawler access