How to measure AI search visibility
By the Motionexa Research Desk · last verified 2026-06-10
You can't manage AI visibility from analytics — ~70.6% of AI-driven visits arrive with no referrer (Digital Bloom, 2026). Measure it like field science: freeze a 40+ prompt panel across five buyer intents, run it monthly in fresh sessions on ChatGPT, Perplexity, AI Overviews, Gemini and Claude, score position-weighted share of voice, log framing adjectives and cited sources, and report the month-over-month delta on identical prompts. Platforms like Profound/Peec/Otterly automate cadence once the number goes to executives; a disciplined spreadsheet is sufficient to start.
The measurement problem
AI answers leave almost no exhaust. Roughly 70.6% of AI-assistant-driven visits arrive with no referrer and hide inside "Direct" traffic [1], and the answers themselves are personalized, probabilistic and re-generated on every ask. So you cannot manage AI visibility from analytics. You measure it the way field science measures anything unstable: fixed instrument, repeated observation, logged conditions. Here is the exact protocol behind our audits — free to copy.
Step 1 — Build the prompt set (the instrument)
40+ prompts across five intents, phrased as buyers phrase them, frozen once written:
| Intent | Share | Examples |
|---|---|---|
| Category shortlist | ~35% | "best [category] for [segment]", "top [category] tools 2026" |
| Competitor displacement | ~20% | "alternatives to [rival]", "[rival] vs [rival2], which is better?" |
| Use-case | ~20% | "how do I [job-to-be-done]", "[problem] solution for [stack]" |
| Commercial detail | ~15% | "[category] pricing comparison", "cheapest [category] with [feature]" |
| Reputation | ~10% | "is [brand] legit?", "[brand] reviews", "[brand] security record" |
Step 2 — Run cadence and logging (the observation)
Each prompt runs on each engine — ChatGPT (browsing on), Perplexity, AI Overviews, Gemini, Claude — in a fresh session, monthly, on a fixed week. Every run is logged with: date, engine, model version, full answer text, brands named, position of each, links cited, and one-line sentiment. One run is an anecdote; the panel over months is data.
Step 3 — Score share of voice (the metric)
Each run is scored on a position-weighted scale — a first-place recommendation counts for more than a list mention, which counts for more than a footnote citation. (The exact weighting is part of the paid methodology; any consistent scale works for your own tracking.) Share of voice = total points ÷ total runs, reported overall and per engine and per intent. Track three companions: coverage (% of prompts where you appear at all), framing (the adjectives engines attach to you, logged verbatim), and citation source mix (which of your/others' pages the engines actually leaned on).
Step 4 — The delta is the deliverable
Re-run the identical, frozen prompt set. Report month-over-month change per intent cluster — that isolates the effect of work shipped between runs. When we publish results, this is the format: same 40 prompts, same engines, dated transcripts, before-column, after-column. Anything else is screenshot theater.
Manual ledger vs monitoring platforms
Dedicated platforms — Profound, Peec, Otterly, and a fast-growing field [2] — automate daily runs, larger prompt panels, and alerting; sensible once AI share of voice is a number an executive asks about monthly. The manual ledger (a spreadsheet and discipline) costs nothing, teaches you what the engines are doing, and is fully sufficient for a quarterly baseline. We run audits manually with scripted logging because the source-reading is the analysis; we recommend platforms to clients who graduate to continuous monitoring.
What to report upward
Three lines on one page: AI share of voice (the weighted score, trended), coverage on money prompts (the 10 prompts closest to revenue), and assisted evidence — "how did you hear about us?" mentions of AI, plus conversion rate of suspected-AI landing sessions (recall AI-referred visitors converted ~5× organic in 2026 data [3]). Resist inventing precision the channel doesn't offer.
Questions people ask
Q.01 How many prompts do I need for a reliable baseline?
40+ across five intent types is our floor for a defensible read on one category; below ~25 a couple of volatile answers swing your score by double digits. Enterprises tracking many segments run hundreds via monitoring platforms, but a frozen 40-prompt panel run monthly beats a 500-prompt panel run once.
Q.02 Why do I get different answers than my colleague for the same prompt?
Personalization (account history, location, model routing) plus sampling randomness. That's exactly why the protocol uses fresh sessions, multiple engines, a large prompt panel and repeated runs — single-screenshot evidence, positive or negative, is noise.
Q.03 What's a good AI share of voice?
Category-dependent, but working bands from our audits: under 10% = effectively invisible; 10–25% = present but losing most asks; 25–45% = competitive; above 45% = you own the answer and the job becomes defense. The trend matters more than the absolute number.
Q.04 Can I just track AI referral traffic instead?
No — about 70% of AI-driven visits carry no referrer (Digital Bloom, 2026), so analytics structurally under-counts the channel. Referral data is a useful floor, never the measure. Share of voice on a fixed prompt set is the primary KPI; self-reported attribution and landing-page inference are the supporting evidence.
Sources & further reading
- [1] The Digital Bloom, AI traffic referrer analysis, February 2026
- [2] Public product documentation: Profound, Peec, Otterly (AI visibility monitoring platforms), 2025–2026
- [3] Exposure Ninja, AI-referral conversion analysis, March 2026
Want this analysis run on your category? The full audit — 40+ prompts, 5 engines, scorecard, source map, fix worksheet — is a flat $1,200, with the founding-client evidence guarantee.