We sell a measurement. It would be strange not to publish our own. Here is the full method, the prompt list, and the score, updated every month whether it improved or not.
45 prompts, five engines, run twice each: once with live retrieval on, once with it off.
Every row carries how we got it. A composite inherits the lowest confidence of its inputs, and we never average across engines.
Total: named 58 times out of 225 prompt-engine pairs. Twelve of those survive with retrieval switched off. So 79% of what visibility we have is rented, and we are a new company with very little of either. That is the honest starting line, and it is the same number we would report for you.
Every prompt is a real buying question, phrased the way somebody would actually type it. No brand names in the prompt. No leading.
Each one runs twice per engine. First with live web retrieval enabled, then with it disabled so the model can only draw on what it already holds. A brand that appears in both is owned. A brand that appears only with retrieval on is rented, because it is sitting inside a page the model happened to find.
Runs happen on the first working day of each month, from a clean session with no history. We record the naming position and the cited source for every hit.
We do not de-duplicate favourably, we do not drop bad runs, and we do not re-run a prompt because we did not like the answer.
Ten of our own 45, verbatim. We publish all of ours as proof. Your set is generated for your category and buyers — the generation is a method, not a template we hand out.
45 prompts, five engines, five working days. $750 for the first 50.
Start the audit