- Tracking means re-running the same prompts on a schedule and storing the results. A single check is an anecdote, not a measurement.
- Record five things per answer: brand coverage, domain citations, share of voice, position in the answer, and sentiment.
- Being mentioned and being cited are separate outcomes. In our own study, Nike was named in 93% of answers while nike.com was cited twice.
- Across 50 buyer prompts, ChatGPT, Gemini, and Perplexity named the same brand only 21% of the time. Track each engine separately.
- Weekly is the practical minimum cadence. Roughly half the sources these engines cite change from one day to the next.
What does tracking brand visibility in AI search actually mean?
It means running a fixed set of buyer questions through AI engines on a repeating schedule, and recording for every answer whether your brand appeared, how it appeared, and who appeared instead of you.
That is the whole discipline. It works the same way rank tracking does: the same prompts, asked identically, at a regular interval, with every result stored so you can compare this week against last month.
What you are measuring is AI brand visibility — how often and how favourably generative engines name your brand when buyers ask them for a recommendation. It is not a single score you climb. It is a set of five numbers per engine, watched over time.
Opening ChatGPT once and typing your category name is not tracking. It tells you what one answer looked like at one moment, which is useful as a sanity check and useless as a measurement.
The reason this is worth the effort: the audience is now large. Pew Research found 34% of US adults had used ChatGPT by mid-2025, roughly double the share two years earlier, with product research among the most common uses. Your buyers are asking these engines which tool to pick, and there is no page two to fall back on.
What is the difference between a mention and a citation?
A mention is your brand name appearing in the text of an answer. A citation is the engine linking to a page on your domain as the source behind that answer. They are different outcomes and you need to record them in separate columns.
The gap between them can be enormous. In our own 11-day study of 165 AI answers, Nike was mentioned in 93% of them while nike.com was cited as a source just twice. The engines knew the brand perfectly well; they simply reached for YouTube, RunRepeat, and Reddit to back up what they said about it.
That means high mentions with near-zero citations is a normal, common state — not a bug in your tracking. It tells you the engines describe you using other people's pages, which is a different problem from being invisible.
Which metrics should you track?
Five metrics cover the whole picture. Each answers a question the others cannot.
| Metric | What it answers | What a healthy trend looks like |
|---|---|---|
| Brand coverage | In what percentage of tracked prompts is my brand mentioned? | Rising over time, ahead of your closest competitors |
| Domain citations | How often does an engine link my pages as a source? | Engines citing your own content directly, not just naming you |
| Share of voice | Of all brand mentions across your prompts, what share is yours? | Growing share within your defined competitor set |
| Average position | When mentioned, am I named first or fifth? | Position 1 to 3 — early mentions carry the recommendation weight |
| Sentiment | Is the engine describing my brand positively? | Net-positive framing, no recurring negative claims |
Brand coverage is your top-of-funnel number: the share of tracked prompts where the engine names you at all. Share of voice narrows that to a competitive read — your mentions as a percentage of all brand mentions across the same prompts, measured against a fixed competitor set. Coverage can rise while share of voice falls, if your competitors are growing faster than you are.
Average position matters because the first brand named in an answer carries the recommendation weight. Being fifth in a list of six is closer to invisible than the coverage number suggests.
Sentiment is the one most people skip, and it does not follow from the others. Tracking four sneaker brands across 258 answers, we found the most-mentioned brand was also the least warmly described — 3,646 mentions but positive tone only 64% of the time, against 82% for a brand named far less often. Volume and warmth are separate axes.
Note: track all five per engine, not blended into one site-wide average. A blended number hides the case that matters most — strong in one engine, absent in another.
How often should you scan, and why isn't one check enough?
Weekly is the practical minimum. Daily is better, because AI answers move far more than most people expect.
We ran the same five buyer questions through ChatGPT, Gemini, and Perplexity every day for 13 consecutive days and compared each day's cited sources against the day before. 47.5% of cited sources changed from one day to the next. Just 4.3% of domains held their place on 12 or more of the 13 days, and the median cited source lasted two days before disappearing.
That churn is the reason a single check misleads. If half of what you observe is a one-day appearance, then seeing your page cited once tells you almost nothing about whether it will still be there tomorrow — and seeing a competitor cited once does not mean they have won the prompt.
Two practical consequences:
- Judge on rolling averages, not single days. A 7-day or 14-day average of your coverage number is stable enough to act on. A daily figure is not.
- Give changes four to six weeks. Engines typically take a few weeks to surface a newly published, well-indexed page. Log every content action with its date so you can connect a shift in the trend to the work that caused it.
This is also the honest argument for automating the scans. Manual checking survives about two busy weeks, and an interrupted dataset cannot show you a trend.
Which AI engines and surfaces do you need to track separately?
All of them that your buyers use — and there are more of them than "ChatGPT, Gemini, Perplexity" suggests, because Google alone answers questions on three different surfaces.
Engine-level disagreement is the starting point. Across 50 buyer-intent prompts run through all three major engines, all three named the same brand only 21% of the time, and 53% of brand mentions came from just one engine. Visibility does not transfer. A tool can read as the obvious category leader inside ChatGPT and be absent from Gemini for the identical question.
Tone varies by engine too: in our sentiment study, the same four brands scored 87% positive on Gemini but 70% on ChatGPT.
| Surface | What it is | Webmaster-tool visibility | How to track it |
|---|---|---|---|
| ChatGPT | Assistant blending training data with live browsing | None | Prompt scans only |
| Perplexity | Searches the live web on every query, shows inline sources | None | Prompt scans; sources are visible in the answer |
| Gemini | Google's standalone assistant | None | Prompt scans |
| Google AI Overviews | AI answer above the normal Google results | Partial — impressions are folded into Search Console's Web totals, not broken out | Prompt scans plus Search Console for the query side |
| Google AI Mode | Google's conversational search interface | None broken out | Prompt scans |
| Microsoft Copilot | Assistant running on Bing's index | Partial — Bing Webmaster Tools has an AI Performance section in beta | Bing Webmaster Tools plus prompt scans |
The practical point buried in that table: Gemini and Google AI Overviews are not the same surface. They draw on Google's index but present answers differently, and they do not reliably name the same brands for the same question. If AI Overviews matter to your category, scan them as their own line item rather than assuming your Gemini number covers them.
You do not have to cover all six from day one. Start with the two or three your buyers actually use, and add surfaces as you find evidence they matter.
Which prompts should you actually track?
Start with real buyer language, not category keywords. "Best project management tool for a 5-person remote team" tells you far more than "project management software," because that is closer to what someone actually types into ChatGPT right before they buy.
Pick 10 to 20 fixed prompts and keep the wording identical every time you scan. Changing phrasing mid-quarter breaks your ability to compare this week against last month.
Cover a mix of shapes, not just one:
- Best-of prompts. "Best [category] for [use case]" — the most common shape, and the one most competitors are also targeting.
- Comparison prompts. "[Competitor] vs [alternative]" or "[competitor] alternatives" — these surface whether you get named as the alternative.
- Problem-first prompts. "How do I [solve the problem your product solves]" — catches buyers earlier, before they've named a category.
- Feature or constraint prompts. "[Category] tool that does [specific feature]" — narrower, but often where a smaller brand can actually win a mention.
One wrinkle worth knowing before you finalize the list: the prompt you type is rarely the only question the engine answers. Modern AI search uses query fan-out — it silently splits your question into several related sub-searches, runs them at once, and merges the results into one reply. Google has confirmed it builds this into AI Mode and Gemini-powered search.
That matters for prompt selection because a page that only answers your exact tracked prompt can still miss the hidden sub-questions the engine is really checking behind it. Before locking in your list, run your top candidates through our free Query Fan-Out Generator to see the sub-queries an engine would plausibly run behind each one, then check whether your page actually answers those too.
How do you set up a tracking workflow you'll actually keep running?
Here is a setup that stays manageable:
- Lock in your prompt list. Use the mix of prompt shapes above, write them as full questions since that's how people talk to an assistant, and don't edit the wording once you start scanning.
- Define your competitor set. Share of voice only means something against the brands you actually lose deals to. Pick 3 to 5 real competitors and keep the set fixed — changing it mid-quarter breaks the comparison.
- Run a manual baseline first. Before automating anything, work through your prompts by hand once so you know what the answers actually look like. Our walkthroughs for checking ChatGPT and checking Perplexity cover the mechanics, including starting a fresh conversation each time so earlier answers don't contaminate the next one.
- Then move to a scheduled scan. Daily if you can automate it, weekly at minimum if you can't.
- Record per-engine results in fixed columns. For each prompt and engine: mentioned (yes/no), position in the answer, cited (yes/no, and which URL), sentiment, and which competitors appeared when you didn't.
- Review weekly, act monthly. Find the prompts where competitors appear consistently and you don't, fix those first, and give the change four to six weeks before judging it.
If you want the citation layer in more depth — which specific URLs get cited, how to set up GA4 and Bing Webmaster Tools for it, and how the dedicated tools compare — that is covered separately in our guide to tracking AI search citations and sources.
Lead Rescue runs this workflow automatically: daily scans of your tracked prompts across ChatGPT, Gemini, and Perplexity, with coverage, citations, share of voice, position, and sentiment in plain language. No SEO knowledge needed to read the results.
What do you do when the numbers look bad?
Read the shape of the problem first. Each pattern points at a different fix, and doing them in the wrong order wastes months.
- Absent from a prompt entirely. No page of yours answers that question well enough to be retrieved. Publish one. Optimising structure on a page that doesn't exist is not a strategy.
- Mentioned but never cited. The engines know your name and use other people's pages to describe you. Make your own pages easier to quote — answer-first openings, question-shaped headings, FAQ schema — and work on the third-party sources they reach for instead.
- A competitor wins the same prompts every time. Look at which sources cite them and get listed in those same places. Our breakdown of why ChatGPT recommends your competitor covers the usual causes.
- Visible in one engine, invisible in another. Expected, not alarming. Diagnose the missing engine on its own terms rather than assuming a site-wide problem.
- Mentioned with poor sentiment. This is a content problem, not a visibility one, and more mentions won't fix it. Find the source the engine is echoing.
Before any of that, confirm your robots.txt actually allows AI crawlers. A blocked crawler makes every other fix pointless.
What can't AI visibility tracking tell you?
Being straight about the limits makes the numbers more useful, not less.
- It is a sample, not a census. You are tracking 15 prompts out of the thousands your buyers might type. The trend is real; the absolute percentage is specific to your prompt set and not comparable to anyone else's.
- Google Search Console can't see ChatGPT or Perplexity. It reports on Google Search only. Google's Search Console help documentation confirms that clicks and impressions from AI Overviews follow the same counting rules as regular results, but it does not offer a way to isolate AI Overview traffic from the rest of your Web search data — you cannot cleanly separate "seen in an AI Overview" from "ranked normally." It is a useful input, not an AI visibility tool.
- Mentions are not traffic. Many AI answers name a brand without linking to it, so visibility gains don't map cleanly onto sessions in your analytics.
- It won't tell you why. Tracking shows you that a competitor took a prompt; working out the reason still means reading the actual answers and the sources behind them.
- Personalisation and location shift results. Answers can vary by user context, so treat your numbers as directional rather than absolute.
None of that undermines the exercise. It just means the right question is "is this trending up against the same competitors on the same prompts?" rather than "what is my score?"
When you're ready to move from measuring to improving, start with our guide on how to get your brand mentioned by ChatGPT.
