AnswerRange
Preprint · 2026–09–10

Where the Questions Come From: Prompt Provenance as a Bound on Generative-Search Visibility Measurement

Muhammad Arshad

DOI 10.5281/zenodo.22694993 · CC BY 4.0

Download PDF View on Zenodo

Abstract

Commercial products report how often generative search engines cite a business. Every such number is conditional on a question set, and that question set is a free parameter: the same brand, measured on the same platforms in the same week, can be assigned very different visibility depending on where its questions came from. We name this parameter prompt provenance and give it a three-class taxonomy — natively authored, machine-generated, and observed search queries. We report two experiments. In the first, a single brand is measured twice on the same five generative-search surfaces with the same replicate structure, differing only in corpus provenance, in four language markets (English, German, Spanish, French; 864 observations). The machine-generated corpus returns the higher citation rate in all four, but the size of the gap varies more than thirteenfold. In the second, a leather-goods retailer measured against a mismatched category corpus returns 0 of 108 — a precise-looking zero that carries no information about the business. We then show that in a census of 16 commercial products, none discloses the provenance of its prompts.

Findings

  1. Corpus provenance moved the headline figure by 24.1 points on one brand in one market (56.5% against 32.4%, two-sided Fisher exact p = 5.9 × 10⁻⁴) — larger than most of the period-over-period changes these products are sold to detect.
  2. The direction held in all four language markets but the magnitude did not: the gap ranged from 1.9 to 24.1 points, a thirteenfold spread within a single business. A fixed bias could be corrected by a footnote; one that varies by an order of magnitude can only be disclosed.
  3. A mismatched corpus produces a measurement that is arithmetically correct and substantively empty, and the failure is invisible in the statistics: an observed zero carries a narrower interval than any interior value, so the least informative measurement looks the most precise.
  4. Observed-query provenance has an eligibility floor. Two further properties in the same category could not support it at all, which means the class is unavailable to exactly the small businesses most likely to be buying a visibility measurement.

Data and code

The full observation set, both prompt corpora and the measurement harness are released under CC BY 4.0 with the record at 10.5281/zenodo.22694993, so every figure above can be recomputed rather than taken on trust.

Cite this work
Arshad, M. (2026). Where the Questions Come From: Prompt Provenance as a Bound on Generative-Search Visibility Measurement. Zenodo. https://doi.org/10.5281/zenodo.22694993

← All papers