- Free GEO tracking works: a spreadsheet and the AI engines themselves are enough to measure real progress.
- Consistency matters more than frequency. The same 10-15 prompts, recorded the same way, every two weeks beats ad hoc checks whenever you remember.
- Track five things per session: brand mentioned (yes/no), position in answer, URL cited, competitor names, and sentiment.
- Manual tracking is accurate but time-consuming. It becomes impractical around 15+ prompts across 3 engines every two weeks.
What does measuring GEO progress actually mean?
Measuring GEO progress means tracking whether your brand is appearing more often in AI answers over time, not just checking once and hoping. A single check tells you what one answer looked like at one moment.
That is not a measurement. Progress requires comparing the same prompts, asked the same way, across multiple sessions weeks apart.
Without that consistency, you have no way to tell whether a change you made actually worked.
GEO progress is measured through the same five metrics that matter in any AI visibility strategy: brand coverage (how often you appear), domain citations (whether your pages are linked), share of voice (your mentions vs competitors), average position (where in the answer you appear), and sentiment (how you are described).
A 2023 Princeton study on generative engine optimisation confirmed that visibility in AI answers is measurable and responds to structured content changes, with improvements of up to 40% observed from specific optimisations (Aggarwal et al., Princeton, 2023). Those improvements only become visible when you are tracking consistently.
The good news is that you do not need a paid tool to start. The bad news is that doing this rigorously by hand is genuinely time-consuming once your prompt set grows past 15 queries across three engines.
What should you track in each session?
For each prompt you run, record five things: whether your brand was mentioned, the position your brand occupied in the answer (first, second, or later), whether a URL from your domain was cited as a source, which competitor names appeared, and whether your brand was described positively, neutrally, or negatively. These five data points, collected consistently, give you a complete picture of your AI visibility without needing any external tool.
| What to record | What it tells you | How to capture it |
|---|---|---|
| Brand mentioned (yes/no) | Coverage — your baseline visibility | 1 or 0 in a spreadsheet cell |
| Position in answer | Whether you lead or trail recommendations | 1, 2, 3, or "not mentioned" |
| URL cited | Whether engines trust your pages as a source | Paste the cited URL or "none" |
| Competitors named | Who is winning the prompts you are losing | Comma-separated brand names |
| Sentiment | How your brand is being described | Positive / neutral / negative |
Rand Fishkin of SparkToro, who tracks AI answer patterns across thousands of queries, has noted that position in the answer matters almost as much as whether you appear at all. Brands named first in an AI recommendation carry the most weight with readers, so tracking position separately from coverage reveals whether you are improving in quality as well as frequency.
How to build a free prompt tracking spreadsheet
A simple Google Sheets or Excel spreadsheet is all you need. Set it up once and it becomes a running record of your GEO progress.
The key is a structure that makes each session take minutes to fill in, not hours, and that makes trends visible at a glance without any formulas or dashboards.
- Create one sheet per engine. Three tabs: ChatGPT, Gemini, Perplexity. Each tab holds the same structure so you can compare across engines easily.
- Rows = prompts, columns = sessions. List your 10-15 prompts down the left column. Each time you run a check, add a new date column to the right. This gives you a visual timeline at a glance.
- Use a compact encoding per cell. For each prompt-session cell, record: M/P/N/X (Mentioned/Position/No-mention/X for error), a citation URL if present, and a one-letter sentiment (P/N/U for positive, negative, unclear). Example: "M2, example.com/page, P" means mentioned in position 2, that URL was cited, described positively.
- Add a summary row at the top. Count the number of M results per session column. That single row, tracked over time, shows your brand coverage trend without any formulas beyond a basic COUNTIF.
- Colour-code competitor appearances. In a separate "competitors" column per cell, mark any competitor names. Highlight in yellow when a specific competitor appears. Red patterns in that column show you exactly where you are losing.
Pew Research found 34% of US adults had used ChatGPT by 2025, with product research among the top use cases. The buyers filling those answers are real.
Having even a simple spreadsheet record lets you see whether your content work is moving you in front of them.
How often should you run manual checks?
Every two weeks is the right cadence for most founders doing this manually. Weekly is better but unsustainable alongside a full-time build schedule.
Monthly is the minimum that still gives you a meaningful trend. What you want to avoid is checking once after a content change, seeing no movement, and concluding the change did not work.
AI answers update on different timescales: Perplexity refreshes fastest (days to a week for newly indexed content), ChatGPT's browsing mode is slower and less predictable, and Gemini follows Google's crawl cycle.
The practical rule: make a change (new page, new listing, updated description), wait three weeks, then run a full session. That three-week window gives Perplexity time to index and re-answer, ChatGPT time to browse the updated page, and Gemini time to reflect the Google index update.
Checking sooner than that is how you incorrectly conclude a change did not work when it just has not propagated yet.
For context on what each engine's update cycle looks like, our guide on how to check if AI engines mention your brand covers the differences in retrieval speed across ChatGPT, Gemini, and Perplexity.
What counts as real progress vs random noise?
Real progress is a sustained directional change across at least three consecutive sessions. A single session where you appear in 8 out of 10 prompts after appearing in 5 the session before is not real progress.
It might be. But one session is noise.
Three sessions in a row showing an upward trend, on the same prompts, is signal. The Princeton GEO research found meaningful, reproducible visibility improvements of up to 40% from content optimisations, but those improvements only become visible with consistent measurement over weeks, not days.
Common sources of noise to recognise and filter out:
- Phrasing variation. A slightly different prompt wording can change which brands appear. If your coverage jumps, check whether you inadvertently asked the prompt differently than last time.
- Session-to-session engine variability. AI answers fluctuate. Two identical prompts run one hour apart can return different results. Track trends over sessions, not individual answers.
- New competitor entries. A competitor launching a new product or getting press coverage can briefly dominate answers across your prompts. A sudden drop in your share of voice may be their spike, not your decline.
- Engine updates. Perplexity, ChatGPT, and Gemini all update their models periodically. A broad shift in results across all your prompts at once often signals a model update, not a change you caused.
When does manual tracking stop being enough?
Manual tracking stops being enough when the time it takes to run sessions consistently starts crowding out the content work that actually moves the numbers. For most founders, that threshold arrives around 15 prompts across three engines, run fortnightly.
That is 45 individual checks per session, plus recording and analysis time. Done properly, it takes 90-120 minutes every two weeks, and it has to be done on a fixed schedule to remain reliable.
The other signal that you have outgrown manual tracking: you start skipping sessions because the friction is too high. A month of missing data breaks your trend line and leaves you with no baseline to compare against when you want to evaluate a specific change.
At that point, the cost of the manual method exceeds the cost of a lightweight tool.
The goal of manual tracking is not to do it forever. It is to build enough of a baseline, over 8-12 weeks, that you understand your own visibility well enough to know what to fix.
Once you have that foundation, our guide on how to track brand visibility in AI search covers what a more systematic measurement approach looks like.
