Preprint · 2026–09–06
Cited or Invisible? A Multilingual, Multi-Platform Audit of Brand Citation in Generative Search
Muhammad Arshad
DOI 10.5281/zenodo.22491737 · CC BY 4.0
Abstract
Generative search engines increasingly answer commercial questions directly, citing a handful of sources rather than returning a ranked list. Which businesses those systems cite is becoming consequential for market access, yet the question is mostly discussed through vendor marketing rather than measurement. We report a pre-registered-style audit of brand citation across four generative answer surfaces (Google AI Overviews, Gemini with search grounding, Perplexity, and ChatGPT with web search), six European languages, and five commercial verticals. Prompts are authored natively in each language rather than translated, because translation measures the translator and not the market. Each prompt is issued three times per platform, giving 4,320 observations with no failed calls.
Findings
- The platform a question is asked on dominates the language it is asked in: citation rates for the same brands over the same questions range from 16.2% on Google AI Overviews to 91.2% on Perplexity, with non-overlapping confidence intervals, while the range across six languages is 41.0% to 61.1% with intervals that overlap almost throughout.
- Prompt selection readily manufactures a language gap. On one platform, a single-pass study of eight prompts per language reports a difference of at least 20 points 16.4% of the time when the corpus-wide difference is 8.3 points; on another it does so 12.5% of the time when the true difference is exactly zero. Our own pilot produced a 38-point gap by this mechanism.
- Repeated identical queries disagree with themselves at rates differing by an order of magnitude across platforms (0.3% to 8.1% of prompt-platform cells non-unanimous), which bounds the precision of any single-shot visibility metric.
- Citations are dispersed rather than concentrated: 5,755 distinct domains across 28,371 citations, with the ten most-cited accounting for 7.7% of the total.
Data and code
The full observation set, both prompt corpora and the measurement
harness are released under CC BY 4.0 with the record at
10.5281/zenodo.22491737, so every figure above can be
recomputed rather than taken on trust.
Cite this work
Arshad, M. (2026). Cited or Invisible? A Multilingual, Multi-Platform Audit of Brand Citation in Generative Search. Zenodo. https://doi.org/10.5281/zenodo.22491737
← All papers