Measuring AI visibility: how to track ChatGPT, Perplexity and AI Overviews properly
How to make AI visibility measurable: which metrics hold up, what a reliable prompt set looks like, why a “ranking in ChatGPT” doesn't exist – with studies, keyword data and a monthly workflow.
Short answer: AI visibility measures how often and how accurately your brand shows up in AI answers – not at which “position”. Three metrics hold up: mention rate across a fixed prompt set, citation rate (is your URL cited or is only the brand named?), and your share of all brand mentions compared to competitors. Single screenshots are not measurement, because the same question returns a different brand list on almost every run. Repeat, document and measure across several systems, and you still get reliable trends.
The sentence I hear most often on this topic: “I asked ChatGPT and I wasn't in there.” That's a start, but it isn't measurement. It tells you about as much as one Google search in an incognito window – with considerably more randomness.

Why a “ranking in ChatGPT” is not a metric
In January 2026, Rand Fishkin (SparkToro) and Patrick O'Donnell (Gumshoe.ai) published the cleanest numbers on this so far. 600 volunteers ran 12 prompts through ChatGPT, Claude and Google AI (Overviews or AI Mode) a combined 2,961 times, each prompt 60 to 100 times per system. The result, from their write-up:
- The exact same brand list reappeared in fewer than 1 in 100 runs.
- The same list in the same order in fewer than 1 in 1,000 runs.
- Even the number of brands mentioned kept shifting.
Fishkin's conclusion is inconvenient for a lot of tool marketing: anyone selling you a “ranking position in AI” is selling a number that doesn't exist. At the same time, the study points to the way out: frequency is stable. In tight categories, leading providers appeared in nearly every answer – one US cancer center showed up in 69 of 71 ChatGPT answers, roughly 97 % visibility, without ever reliably being “first”.
Hence the ground rule for everything below: you measure probabilities, not places. And probabilities need repetition.
Demand picture for this topic (DataForSEO, July 2026)
Before you invest in tools, look at what demand actually looks like. I pulled the clusters through the DataForSEO APIs – Google volume for Germany plus the AI keyword volume from the AI Optimization module:
| Search term | Google volume/month (DE) | 3-month trend | Intent |
|---|---|---|---|
| generative engine optimization | 1,300 | −33 % | informational |
| chatgpt seo | 480 | flat | commercial |
| ki-sichtbarkeit | 140 | flat | informational |
| ai visibility | 140 | −15 % | informational |
| ki-sichtbarkeit messen | 70 | flat | informational |
| ai visibility tool | 70 | +26 % | commercial |
| geo monitoring | 30 | +22 % | commercial |
| llm monitoring | 20 | +17 % | informational |
Two things stand out. First, tool and monitoring terms are growing, while the umbrella term “generative engine optimization” is coming down after its hype spring. The market is moving from “what is this?” to “how do I measure it?”. Second, the AI search volume – demand inside LLM prompts, which DataForSEO reports separately – sits at zero to three prompts per month for nearly all of these terms. That doesn't mean nobody asks AI about them. It means the data basis for prompt demand is still thin, so don't build your planning on it. Classic keyword volume remains the better proxy; real customer questions are the best one.
The SERP for “ai visibility” also shows who currently owns the space: Google displays an AI Overview citing eight sources – among them ahrefs.com, seranking.com, amplitude.com, llmpulse.ai and two German agency sites. Five of those eight also sit in the organic top 10. The takeaway: in this topic area, classic rankings and AI citations still correlate strongly. If you're missing organically, you're usually missing from the answer too.
What measurement is worth commercially
Measuring costs time. The justification is the quality of this channel.
Similarweb reports a 7.1 % conversion rate for ChatGPT referrals (clickstream panel, April–May 2026) – only paid search is higher at 7.8 %, with classic organic and direct behind it (Gen AI Stats 2026). Semrush's traffic study points the same way: a visitor from an AI source is on average 4.4 times as valuable as an organic visitor, measured by conversion rate.
A second Semrush finding matters even more in practice: pages ChatGPT cites rank in classic search at position 21 or lower almost 90 % of the time. Citation readiness and ranking strength are not the same thing – which means real opportunities for content that would never reach page one.
The reverse direction is documented as well: according to SISTRIX analysis across 100 million keywords (February 2026), AI Overviews appear for around 20 % of German keywords, and where they appear, the click-through rate on position 1 drops from roughly 27 % to 11 %. Visibility is shifting whether you measure it or not.

The four metrics that hold up
| Metric | What it answers | How you collect it | Limit |
|---|---|---|---|
| Mention rate (visibility %) | In what share of answers does the brand appear? | prompt set × repetitions, count mentions | says nothing about links |
| Citation rate | Is your URL cited as a source? | log source chips/links per answer | highly engine-dependent |
| Share of answer | What's your share versus the competitors named? | own mentions ÷ all brand mentions in the set | needs a clean competitor list |
| Accuracy of portrayal | Is your offer described correctly and positively? | read and rate answers on a sample basis | manual, not automatable |
The fourth metric is the one most often forgotten, and frequently the most important. A mention with the wrong value proposition, an outdated price or a mixed-up audience does more harm than good. Being mentioned isn't the goal – being mentioned correctly is.
Keep mentions and citations apart. Perplexity is strongly source-driven; ChatGPT names brands from model knowledge without a link. If you count only citations, you systematically underestimate your ChatGPT visibility; if you count only mentions, you overestimate your influence on the source landscape. Measure both, separately.
A prompt set that actually says something
The prompt set is the denominator of every measurement. If it wobbles, everything wobbles.
Size: 30 to 50 questions is realistic and meaningful for freelancers and SMBs. Below 20 it turns into noise; above 100 nobody maintains it.
Composition by funnel stage:
- Awareness – problem questions without vendor context: “Why is my website losing traffic despite good rankings?”
- Consideration – criteria questions: “What should I look for in an SEO agency for trade businesses?”
- Decision – vendor questions with intent: “Who offers AI search consulting for SMBs in Germany?”
- Brand – direct brand questions: “What does schoettler.io do?” (mainly tests accuracy of portrayal)
- Local, where relevant – “Who does SEO in Bad Oeynhausen?”
Repetitions: The SparkToro numbers show single runs are worthless. A workable compromise: run your five most important prompts five times per cycle and the rest once. You see the spread where it matters without multiplying the effort.
Consistency: wording, order, language, region and time window stay the same. No toggling “AI mode” on and off. Strip personalisation where you can (log out or use a separate profile), otherwise you're measuring your own history.

How many systems do you need?
Here the available data openly contradicts itself – worth knowing before you trust a single number.
One widely quoted audit puts the overlap of domains cited by ChatGPT and Perplexity at just 11 %. The German-language report State of AI Search 2026, covering 12,000 prompts across 150 brands, instead finds 64 % overlap between ChatGPT and Perplexity, 71 % between ChatGPT and Claude, and 81 % between AI Overviews and Gemini. The spread comes from industries, prompt types and measurement windows – not from “wrong” data.
In practice: the core of the work is cross-system (substance, structure, authority), but measurement has to happen per system. Three are enough to start:
| Engine | Why measure it | Particularity |
|---|---|---|
| Google AI Overviews | largest reach in the DACH region | tightly coupled to index and rankings |
| ChatGPT | largest chatbot usage | often names brands without a source |
| Perplexity | most transparent sourcing | ideal for spotting source gaps |
Gemini, Claude and Copilot come in when your audience uses them – not out of a need for completeness.
The monthly workflow in six steps
Here's a cycle that fits into two to three hours a month:
- Run the prompt set – same order, same systems, same time window. Save answers as text and screenshot.
- Log mentions – per answer: own brand mentioned (yes/no), own URL cited (yes/no), competitors named, position in the text (top/middle/bottom).
- Calculate metrics – mention rate, citation rate, share of answer per engine and per funnel stage.
- Analyse source gaps – which third-party pages get cited for your core questions? Those are your content and PR targets.
- Derive two actions – not twenty. For example: one core question gets a citable answer block, one service page gets clear criteria and price ranges.
- Re-measure after 30 days – same conditions, document the change.
What that realistically produces is visible in a practice analysis by getSichtbar across four GEO projects: the citation rate for their own URLs rose from 9.6 % to 32.8 % (AI Visibility Benchmark 2026, based on 13,112 GEO checks). No guarantee, but a solid signal: source work has measurable effect – over months, not overnight.

Tools: what they do and don't do
The tool landscape in 2026 is crowded, with prices from roughly €50 a month into four figures. What they deliver:
- Automated prompt runs across several engines – the actual time saver.
- Competitive comparison with an open brand list.
- History and reporting, so you see trends instead of snapshots.
What they don't deliver:
- No real “rankings”. Anyone promising positions has either missed the volatility or is ignoring it.
- No judgement on accuracy. Whether your offer is described correctly is something you read yourself.
- No completeness. Every tool picks engines, regions and a sampling method. Ask before buying: how many runs per prompt? Which engines in which mode? Is personalisation excluded? How is the denominator for share of answer built?
To start, a spreadsheet is enough. Honestly: a sheet with prompt, date, engine, mention, citation and competitors gives you more insight after two cycles than a tool subscription nobody reviews. A tool pays off once your set passes 30 prompts or you're looking after several brands.
Making AI traffic visible in analytics
Alongside answer measurement, watch what actually lands on the site. In GA4 or your analytics tool, filter referrers such as chatgpt.com, perplexity.ai, gemini.google.com and copilot.microsoft.com into their own segment.
Three limitations you need to know:
- A large share of AI clicks arrives without a referrer and lands under “direct”. Your measured AI visibility is a lower bound.
- AI Overviews barely produce distinguishable traffic – in Search Console they look like normal Google clicks. The symptom of an Overview effect: impressions steady, clicks falling.
- Mentions without clicks still work. They show up later in brand searches and direct visits. Watch brand traffic as an indirect indicator.
Common mistakes
Mistake: Ask once, screenshot, change strategy. Better: Repeat, count, act only on a stable pattern.
Mistake: Measuring ChatGPT only. Better: At least AI Overviews, ChatGPT and Perplexity – and evaluate each engine separately.
Mistake: “Improving” prompt wording between measurements. Better: Freeze the set. Document changes and treat them as a new time series.
Mistake: Counting mentions only and celebrating content. Better: Track citation rate and accuracy too – that's where the work is.
Mistake: Measuring without consequence. Better: Two actions per cycle, executed properly.

Conclusion: measurable instead of speculative
AI visibility is measurable – just not the way classic rankings taught us. Positions are noise, frequencies are signal. A fixed prompt set, three engines, four metrics and a monthly cycle are enough to turn gut feeling into decisions you can defend. Everything else – tools, automation, dashboards – is convenience, not a substitute.
If you want to know where you stand today: in a free intro call I'll build a first prompt set with you and measure the baseline. The fundamentals are in my article AI Search & GEO 2026, the ongoing work in the service AI Search & GEO. If you need a broader assessment first, the marketing audit is the faster entry point.
FAQ on AI visibility
How often should I measure? Monthly is enough for freelancers and SMBs. More important than frequency is keeping the prompt set, engines and conditions constant. Weekly only pays off when you're actively working on content and need fast feedback.
Do I need a paid AI visibility tool? Not to start. Up to around 30 prompts, a spreadsheet with documented answers is sufficient and often more informative, because you actually read the answers. Beyond several brands or larger prompt sets, a tool saves real hours.
Why do I see different providers for the same question every time? Because generative systems don't answer deterministically. The SparkToro/Gumshoe study found the same brand list in fewer than 1 in 100 runs. That's exactly why you measure frequency across many runs instead of positions in a single answer.
Does a mention without a link count at all? Yes. It shapes the choice someone makes and shows up later in brand searches and direct visits. It's still weaker for steering than a citation, because you have less leverage on the source landscape – so track both metrics separately.