Muhammad Arshad
DOI 10.5281/zenodo.22694993 · CC BY 4.0
Commercial products report how often generative search engines cite a business. Every such number is conditional on a question set, and that question set is a free parameter: the same brand, measured on the same platforms in the same week, can be assigned very different visibility depending on where its questions came from. We name this parameter prompt provenance and give it a three-class taxonomy — natively authored, machine-generated, and observed search queries. We report two experiments. In the first, a single brand is measured twice on the same five generative-search surfaces with the same replicate structure, differing only in corpus provenance, in four language markets (English, German, Spanish, French; 864 observations). The machine-generated corpus returns the higher citation rate in all four, but the size of the gap varies more than thirteenfold. In the second, a leather-goods retailer measured against a mismatched category corpus returns 0 of 108 — a precise-looking zero that carries no information about the business. We then show that in a census of 16 commercial products, none discloses the provenance of its prompts.
The full observation set, both prompt corpora and the measurement harness are released under CC BY 4.0 with the record at 10.5281/zenodo.22694993, so every figure above can be recomputed rather than taken on trust.