How we measure AI search
Everything published in Insights comes out of one dataset and one set of rules. This is the rulebook: what a citation is, where the numbers come from, the evidence floors that decide what we can name, what we will never publish, and what we do when we get something wrong.
Behind every figure: 264,296 citations · 55,285 sourced responses · 0 AI models · measured, never estimated
Maintained by the BrandInsightAI Data Desk · last reviewed 26 August 2026
On this page
- How we read what AI searched for
- How desire is measured
- What counts as a citation
- Where the data comes from
- Complete weeks, and why movement uses only those
- The floors that decide what we can name
- What we never publish
- Facts first, language second
- The figure contract
- What this data cannot tell you
- Corrections
How we read what AI searched for
The searches the assistants run on their own behalf, which assistants report them, and what we do and do not claim.
Before an assistant answers, it usually runs its own web searches. Those searches come back with the answer, and we store them exactly as issued. They are the assistant's shortlist: whoever it searched for is the field it then chose from, which is why a brand absent from them was never really a candidate, however good its content is.
Four assistants report their searches — Gemini, Claude, ChatGPT and ChatGPT Web. Perplexity, Google AI Overview, Google AI Mode, Grok, DeepSeek and Meta AI do not. Every share we publish is measured against the four that do, and we say so on the page rather than leaving it implied.
Assistants reword their own searches constantly: run the same prompt on consecutive days and only about a fifth to a third of the wording repeats, even though the underlying question has not moved. So we do not compare the strings. Each search is reduced to its content words, with brand names, years and market names replaced by placeholders, and it is that fingerprint that is compared over time. Each search is also labelled with what it was trying to find out — comparing options, checking price, looking for the best, going straight to a source, and so on — using fixed keyword rules, in every language the programme runs in. A search we cannot place is left as a general question rather than guessed at.
Three limits. These are the assistants' searches, never what a person typed into a search engine, so nothing here is search volume. Searches and the sources an answer cites are stored side by side but not linked, so we can say an answer ran six searches and cited nine pages, never which search produced which page. And where a category edition quotes a search, it is quoted verbatim from a programme category, never reconstructed.
How desire is measured
The three drivers, the sixteen dimensions beneath them, and what an index of 100 means.
Every week, each category in the programme runs a fixed set of unbranded consumer prompts against the assistants. A separate model then reads those answers and scores every brand named in them across 16 perception dimensions, 0–10, using only what the answers say. Nobody is surveyed; nothing is estimated.
The dimensions are grouped into the three Drivers of Desire from the Havas Science of Desire framework, which is the same structure Havas uses to measure desire in people:
- Attraction — standing out and being wanted: quality, innovation, style, boldness, differentiation, leadership.
- Affinity — belonging: lifestyle fit, sustainability, trust, heritage.
- Attachment — converting desire to demand: service, delivery, value, ease, range, reliability.
A brand's driver score is the mean of its scored dimensions in that driver. The index divides that by the mean of every indexed brand across every category in the index and multiplies by 100, so 100 is the average brand anywhere in the index, 115 is fifteen per cent above it, and every brand and every category shares one denominator. A category's own mean, indexed the same way, is how desired the category is. A brand needs at least six scored dimensions in a run to be indexed, and a category needs at least four such brands to publish a board.
Two limits are worth stating plainly. This measures how assistants describe brands, not what people feel — the two can disagree, and that gap is itself interesting. And where we report that a driver “predicts” being named, that is a rank correlation across the brands in one category in one run: a relationship, not a cause.
What counts as a citation
The unit of measurement, and why it is the site rather than the link.
news.bbc.co.uk, www.bbc.co.uk and bbc.co.uk all count as bbc.co.uk.A citation is one domain cited in one AI response. That is the atom every figure on this site is built from. Counting each link separately would let a single long answer with a heavy footnote list outweigh a hundred ordinary ones, so a domain cited repeatedly within one answer still counts once: the question we are answering is which sites do models reach for, not how many footnotes they attach.
Every URL is reduced to its registrable root before it is counted, which means shares here are shares of sites, not of pages or links. Subdomains, tracking parameters and www prefixes all collapse away.
Where one AI product has more than one surface, the surfaces fold together. ChatGPT Web, the web-grounded variant, counts as ChatGPT, because it is browsing mode on the same product rather than a separate model. Splitting them would understate ChatGPT and invent a model nobody chooses to use.
A model only appears in the citation figures if its answers actually carry source links. Today that means Perplexity, ChatGPT, Claude, Gemini, Google AIO and Google AI Mode. Grok, DeepSeek and Meta AI answer without citing anything: they run in the dataset and are measured elsewhere on the platform, but they contribute no citations and can never move a share here.
One piece of small print worth knowing when you read a count ribbon: the responses figure counts sourced responses, answers that returned at least one link. It is the denominator behind "citations per answer" and nothing else.
Where the data comes from
Real analysis runs, aggregated into weekly snapshots that recompute hourly.
The source is the BrandInsightAI macro dataset: large-scale analysis runs that put structured category questions to the models, with each answer's source list captured exactly as it comes back. There is no panel, no scraping of the models' front ends for effect and no sampling step: the population is the set of answers the platform genuinely received.
Those answers are aggregated into weekly buckets: Monday-to-Monday ISO weeks, keyed on domain and model. The in-progress week's bucket is recomputed every hour; finished weeks are already final and are never rewritten. The living leaderboards read those buckets rather than the raw answers, which is why they can be rendered fresh on every request and still be quick.
Being honest about what this dataset is: it is a macro dataset, not a random sample of the questions the world asks an AI. It leans towards the categories and markets under the heaviest measurement. Everything here describes this dataset. Where a finding could plausibly be an artefact of that skew, we say so on the piece rather than leaving you to guess.
Complete weeks, and why movement uses only those
The in-progress week is real data; it is just not a week yet.
A week is complete once it is over. The current week is rewritten hourly as new answers land, so on a Tuesday morning it holds two days of data and by Sunday night it holds seven.
Every week-over-week comparison on this site therefore uses the two most recent complete ISO weeks. Comparing a live week against a finished one would manufacture a collapse every Monday morning and a miraculous recovery every Friday afternoon: the movement would be an artefact of the clock, not of the models.
Rolling windows behave differently and deliberately so. A four-week leaderboard does include the current week, because a share computed across the whole window is not distorted by its last few days being thin: the partial week shrinks both the numerator and the denominator together.
The floors that decide what we can name
Three evidence thresholds. They govern naming, never arithmetic.
One project citing a domain heavily is one corner of the dataset. Two unrelated projects citing it is the minimum evidence that the behaviour belongs to the model rather than to a single measurement brief, and that is the bar a domain has to clear before we will print its name.
Movers carry a second floor because small denominators produce spectacular percentages: a domain going from two citations to six is a 200% rise and means nothing. Requiring the citation floor in both weeks does a second job, a domain that only entered monitoring this week cannot appear as the week's biggest gainer, which is the single most common way a citation leaderboard lies to itself.
The floors decide what can be named. They never decide what is counted.
This distinction is the one worth being pedantic about. Every share is calculated against all citations in scope, including those belonging to domains too thinly evidenced to name. So the percentages on a leaderboard will not add up to 100: the remainder is a long tail we can see perfectly well and simply will not identify. The alternative, dropping the hidden citations out of the denominator, would inflate every visible domain by exactly the amount we were trying to be careful about. A named domain's 3.1% means 3.1% of everything, not 3.1% of what survived the gate.
What we never publish
A closed vocabulary, enforced twice, failing closed.
These pages are built out of a deliberately narrow vocabulary: cited third-party domains, AI model names, market names, closed enumerations such as content types and journey stages, period labels, and numbers. Anything outside that vocabulary is refused rather than redacted: the piece stops instead of being quietly cleaned up.
- Tracked brand and competitor names. No brand tracked on the platform is ever named, nor any of its rivals.
- Project names. Nothing identifies which part of the dataset a figure came from.
- Prompts and queries. No verbatim question, no paraphrase, and no fan-out search string a model generated for itself. Where we report on fan-out behaviour we publish counts and structure only; the strings never leave the database.
- Owned and competitor URLs. Only third-party domains that have cleared the naming floor can appear. A URL outside that vocabulary fails the scan on sight.
- Free-text topics. Freely written category descriptions stay private; only platform-authored enumerations are publishable.
Two independent gates enforce this, and both fail closed. Before any figures reach a language model, they are walked leaf by leaf against the allowed vocabulary, map keys included, because an object keyed by brand name is the classic way a dataset leaks, and anything unrecognised aborts the article. After the prose is written, it is scanned again against a live list of every brand name held on the platform, and every URL it contains is checked against the cited-domain vocabulary. A piece that trips either gate does not get edited into shape; it fails.
That brand list is assembled at scan time and thrown away. It is never stored on an article, never written to a log and never leaves the server.
Facts first, language second
The arithmetic is deterministic. The model only narrates it. A person signs it off.
Every number is computed in SQL or code before a language model is involved, from the same snapshots the living leaderboards render. The figures, the charts and the data tables in an article all come out of that one computation.
The model's job is narrow by design: it writes about the numbers it is handed. It has no database access, no lookup tool, and an instruction to use only the supplied figures. If a sentence would need a number that was not provided, the sentence does not get written, and where synthesis fails or trips a gate, we fall back to a deterministic write-up of the very same numbers.
Re-running the prose never re-runs the arithmetic. An editor can send a draft back to be rewritten, and it is rewritten over the stored facts: the figures in the second version are identical to the first, because they were never recalculated.
Nothing publishes itself. The generator can only produce a draft; a person reads every article and explicitly approves it before it becomes public. Each piece then carries its sample size in a ribbon at the top and a method-and-limitations block at the foot.
Finally, the framing. This is observational data. We can see which sites models cite and how that changes; we do not run controlled experiments, so nothing here establishes cause. Findings are written correlationally on purpose, and where a plausible confound exists we name it rather than letting a clean sentence imply more than the data carries.
The figure contract
Every chart ships with a claim and the numbers behind it.
Every chart on this site is a figure with three parts, and it is never published with fewer.
The drawing is inline SVG with an accessible role, no scripts, no external files, nothing to load and nothing to block. The caption states the claim in words rather than restating the title, so a reader who takes only the caption still takes away something true. The table holds the underlying numbers as real HTML, sitting in a collapsed block beneath the chart and present in the page source whether it is opened or not.
The point of the third part is that a sighted reader, a screen-reader user and a crawler all get the same numbers. If you want to check a figure, quote it or reuse it, expand the table, that is what it is there for.
What this data cannot tell you
The honest edges of the measurement.
A citation is not a visit. We measure what models cite, not what people click. A site can be cited constantly and receive very little traffic from it, and the relationship between the two is not something this dataset can settle.
Models change without announcing it. A step change in a chart can be a retrieval or ranking change inside a model rather than anything a publisher did. We flag abrupt discontinuities where we spot them, but we cannot always tell the two apart from the outside.
Coverage follows the dataset. Categories and markets under heavier measurement are represented far better than those under lighter measurement. Absence from a leaderboard is not evidence that a site is never cited: it may simply sit outside what we observe.
Shares are relative. A domain can lose share in a week when its own citation count grew, because something else grew faster. Where that has happened we say so instead of writing the fall as a decline.
Corrections
We correct in place, and we say what changed.
If something here is wrong, we would rather fix it than defend it. Corrections are made in place: the page keeps its original URL rather than being quietly replaced or taken down, and a dated correction note is added to the page saying what was wrong and what changed. Nothing is removed from the record.
Where the error is in the computation rather than the wording, the affected figures are recomputed and every page that carried them is updated, including any chart's data table, which must always match the drawing above it.
Living leaderboards are a special case. They re-render from the current snapshots on every request, so a fix to the underlying data propagates on its own within the hour and needs no correction note. When a fix changes a definition (a floor, a window, a fold), we note it on this page instead, because a definition applies to every page at once. This page's last review date is at the top for exactly that reason.
If you believe a figure on this site is wrong, tell us through the contact route on our FAQ page, ideally with the page URL and the number you are querying.
Read the data
The rules above are only interesting applied. These are the pages they govern — every figure on each of them is counted under this methodology.