Asking ChatGPT whether it knows you measures nothing
Asking ChatGPT once whether it knows your company measures nothing, because the answer changes the moment you ask again. Parse, which sells an AI visibility measurement tool, replayed 16,143 ChatGPT questions roughly 22 times each between March 26 and April 25, 2026, and published the result on June 25, 2026: two answers to the same question share, on average, only 21% of their cited sources. An AI answer is a dice roll, not a measurement.
What sits behind that measurement is revenue. Ahrefs compared 300,000 keywords between December 2023 and December 2025 and published the result on February 4, 2026: when an AI Overview appears at the top of the results page, the first organic position receives 58% fewer clicks on average. Knowing whether generative AI cites your business stops being a curiosity the day part of your traffic starts flowing through it.
Why do two answers to the same question cite different sites?
A generative engine composes each answer at request time instead of serving a stored page, so the set of sources it settles on shifts from one call to the next. In the Parse study published on June 25, 2026, the average overlap between two answers to the same question is roughly 21% on ChatGPT and roughly 31% on Google AI Overviews. Parse attributes the churn to model sampling noise rather than to the web changing underneath during the measured window. Parse sells a tool in this category, which gives it a stake in the finding, but the study publishes its method, its window and its volumes.
The practical consequence hits anyone testing by hand. You ask the question, ChatGPT cites your site, you conclude things are fine. A colleague asks the identical question an hour later and your site is gone. Neither answer is wrong: they describe the same site through two different draws. A founder deciding on that basis is deciding on a coin flip.
What has to be measured for a number to mean something?
What can be measured is a rate across a set of questions repeated over time, never a single answer. Biencité asks your questions once a week across five generative AI surfaces and counts the share of questions where your site appears in the answer. The output is not a yes or a no. It is a share that moves week to week, and only its slope carries information.
The weekly cadence is a readability decision, not a capacity limit. Measurements taken closer together capture model sampling noise instead of a change on your site, and produce a jittery line that says nothing. One measurement per week gives one point per week, and a slope becomes readable after several points rather than after two.
Why query five AI engines rather than ChatGPT alone?
Being cited by ChatGPT tells you nothing about what Perplexity answers, and the platforms do not behave alike. The Parse study of June 25, 2026 measures two different degrees of instability on the two platforms it observed, roughly 21% source overlap on ChatGPT against roughly 31% on Google AI Overviews. Two engines that do not churn at the same rate cannot stand in for each other on a dashboard.
Biencité therefore queries ChatGPT, Claude, Perplexity, Gemini and Le Chat, and reports the result surface by surface instead of melting it into a single score. Le Chat, built by the French company Mistral, is the surface most English-language tools leave out entirely, and a company selling into Europe can be cited there while staying absent from US-centric answers. The reverse happens just as often.
Which questions should you track, and how many?
The questions worth tracking are the ones your customers type, never your own brand name. A question containing your company name will cite you almost every time and teaches you nothing: it measures an AI's ability to look up a proper noun, not your place in a buying decision. Three criteria make a question list useful.
- Buyer phrasing, exactly as it gets typed: "best invoicing software for a small agency", "how much does a heat pump installation cost", "who does B2B SEO in Manchester".
- Questions you are willing to lose. If your site shows up on every question in week one, the list is too easy and the line will sit flat at the top forever.
- Questions that hold their meaning over time. A question tied to a news cycle means something different a month later, and week-on-week comparison stops working.
On volume, the Starter plan tracks 5 questions per week. The Pro plan raises that to 15 questions per week across the five AI surfaces and adds share of voice against your competitors. For anyone unsure where to start, the tool proposes 15 questions derived from the site's own content, to keep or discard before the first measurement runs.
What do you do when an AI cites a competitor instead of you?
When an AI answers without mentioning you, the useful question is not "why not me" but "who, then". The Biencité tracker lists, question by question, the domains cited in the answers where your site does not appear. Those domains are not an abstract ranking. They are live public pages that answer, today, a question your customers are asking, and you can open them and read what they say that you do not.
The same three or four domains usually recur across an entire question list, and a weekly alert flags it when a new domain starts appearing regularly in your answers. That turnover moves faster than a ranking change in a classic search engine, which is one more reason a single manual check taken in isolation ages badly.
What does this tracking not measure?
Google AI Overviews are not one of the five measured surfaces. They expose no public interface to query, and collecting them would mean scraping Google's results page, which Biencité does not do. Google now publishes its own AI visibility figures instead, and the article "Google just opened its AI visibility report to every site. It counts impressions, not clicks." sets out what that report shows and what it withholds.
Questions are asked from a configured search location, set to France by default for this market, on the platforms whose interface documents that setting: Anthropic, OpenAI and Perplexity. Gemini and Mistral document no equivalent, and Biencité does not send an undocumented parameter, which would be ignored at best and rejected at worst. The limit is worth knowing before reading a report. On those two surfaces, a strongly local question is measured with no guarantee the search runs from the intended country.
Citation tracking finally measures neither your traffic, nor your revenue, nor your technical eligibility. A site that AI crawlers cannot read will appear in no answer at all, and the tracker will report exactly that, week after week, at a monthly price. Eligibility is a different measurement, the scan, described in the article "What an AI visibility scan actually measures, and what it cannot". The order matters: make the site readable first, then measure what the AI engines do with it.
When is citation tracking not worth paying for?
Citation tracking is not worth paying for while nothing can be done with the result. Four situations come up regularly, and in each of them the money belongs somewhere else.
- A site AI crawlers cannot read. The line will sit at zero, and you will have paid for several months to learn what a free scan reports in a minute.
- A business nobody mentions anywhere. Generative AI cross-references sources, and an isolated site does not become citable because it is being measured.
- A quarter with no changes to the site. Measuring without acting produces a chart, not a result: tracking earns its keep through what happens between two measurements.
- A market where AI is not the buying path. If your customers arrive through tenders, direct referral or a closed professional network, your standing inside ChatGPT is not this year's lever.
Citation tracking promises no citation, and no tool honestly can. It reports an observed rate, on your questions, across five surfaces, week after week. That is a modest claim, and it is still far more solid than one question typed into ChatGPT on a Thursday evening.
Frequently asked questions
How long before a citation chart becomes readable?
Several weeks. A weekly measurement produces one point per week, and the variability of AI answers means a single point cannot be read on its own. The slope over one or two months is what becomes usable, not the gap between two consecutive measurements.
Why not measure every day?
Because closely spaced measurements capture model sampling noise rather than a change on your site. The Biencité measurement runs automatically once a week, and that cadence is a choice about how readable the resulting chart is.
Can I run this tracking myself by hand?
You can, as long as you can sustain the number of draws. Tracking 15 questions across 5 AI surfaces once a week means 75 queries every week, recorded the same way each time so the comparison holds. The logging is what costs time, not the asking.
Does tracking guarantee that AI engines will cite me?
No, and no tool can guarantee that. Tracking observes a citation rate on your questions and shows who gets cited in your place. What AI engines cite after that depends on your content, your reputation and how the question is phrased.
- ai citation tracking
- chatgpt
- ai visibility
- geo
- measurement
- perplexity