Written out in full, because a measurement product that will not show its method is asking to be taken on faith.
Each measurement uses a corpus of buying questions for your vertical — the phrasing customers actually use when asking an assistant for a recommendation, not keyword strings. For every non-English market the questions are authored natively by a speaker of that language rather than translated from the English set.
This matters more than it sounds. German furniture buyers search Nachbau and Dutch buyers say namaak; neither is what a dictionary returns for "replica". A translated corpus measures the translation, and quietly reports a market you did not ask about.
Four are measured: ChatGPT, Google AI Overviews, Gemini and Perplexity. All are reached through provider APIs and Google's AI Overview results. These are not identical to the consumer apps — personalisation, session history and app-only retrieval layers differ. We state that rather than bury it, because it bounds what the number means.
Every question is asked three times per platform. This is not redundancy: these systems do not answer identically twice. Our audit found one platform contradicting itself on roughly 40% of identical repeated questions. A single ask would report that coin-flip as fact.
An answer counts as citing you when your registrable domain appears among the sources the platform
attributes. Domains are normalised through a public-suffix rule, so shop.example.co.uk
resolves to example.co.uk and not to the bare suffix. Brand mentions without a link are
not counted — they are real, but they are a different measurement, and conflating the two inflates
the figure.
Citation rates are reported with a 95% confidence interval computed by bootstrap resampling over questions rather than over individual answers. Resampling answers would treat three repeats of one question as three independent facts and report an interval far narrower than the measurement deserves. Clustering on the question is the conservative choice and the correct one.
A movement between two measurements is reported as a change only when the two intervals do not overlap. Everything else is labelled within noise.
Overlapping intervals do not strictly prove that nothing changed, and this rule will therefore stay silent about some real movement. That trade is deliberate: telling a merchant they dropped when they did not is the more expensive mistake, because they will spend money reacting to it.
Both papers are open access with their full datasets — 4,320 observations, both prompt corpora and the measurement code, under CC BY 4.0.