How to Measure Your GEO Progress Without Paid Tools

How to Measure Your GEO Progress Without Paid Tools

Updated 2026-06-14 · 7 min read · by May, Founder of Lead Rescue

You can measure GEO progress without paid tools by running a fixed set of buyer-intent prompts across ChatGPT, Gemini, and Perplexity every two weeks, recording the results in a spreadsheet, and comparing brand coverage and share of voice over time. The method is free, takes about an hour per session, and gives you a reliable trend if you stick to the same prompts and the same recording format each time.

Key Takeaways
  • Free GEO tracking works: a spreadsheet and the AI engines themselves are enough to measure real progress.
  • Consistency matters more than frequency. The same 10-15 prompts, recorded the same way, every two weeks beats ad hoc checks whenever you remember.
  • Track five things per session: brand mentioned (yes/no), position in answer, URL cited, competitor names, and sentiment.
  • Manual tracking is accurate but time-consuming. It becomes impractical around 15+ prompts across 3 engines every two weeks.

What does measuring GEO progress actually mean?

Measuring GEO progress means tracking whether your brand is appearing more often in AI answers over time, not just checking once and hoping. A single check tells you what one answer looked like at one moment.

That is not a measurement. Progress requires comparing the same prompts, asked the same way, across multiple sessions weeks apart.

Without that consistency, you have no way to tell whether a change you made actually worked.

+40% visibility lift observed in AI answers from specific content optimisations — but only visible with consistent tracking. Aggarwal et al., Princeton, 2023

GEO progress is measured through the same five metrics that matter in any AI visibility strategy: brand coverage (how often you appear), domain citations (whether your pages are linked), share of voice (your mentions vs competitors), average position (where in the answer you appear), and sentiment (how you are described).

A 2023 Princeton study on generative engine optimisation confirmed that visibility in AI answers is measurable and responds to structured content changes, with improvements of up to 40% observed from specific optimisations (Aggarwal et al., Princeton, 2023). Those improvements only become visible when you are tracking consistently.

The good news is that you do not need a paid tool to start. The bad news is that doing this rigorously by hand is genuinely time-consuming once your prompt set grows past 15 queries across three engines.

What should you track in each session?

For each prompt you run, record five things: whether your brand was mentioned, the position your brand occupied in the answer (first, second, or later), whether a URL from your domain was cited as a source, which competitor names appeared, and whether your brand was described positively, neutrally, or negatively. These five data points, collected consistently, give you a complete picture of your AI visibility without needing any external tool.

What to recordWhat it tells youHow to capture it
Brand mentioned (yes/no)Coverage — your baseline visibility1 or 0 in a spreadsheet cell
Position in answerWhether you lead or trail recommendations1, 2, 3, or "not mentioned"
URL citedWhether engines trust your pages as a sourcePaste the cited URL or "none"
Competitors namedWho is winning the prompts you are losingComma-separated brand names
SentimentHow your brand is being describedPositive / neutral / negative

Rand Fishkin of SparkToro, who tracks AI answer patterns across thousands of queries, has noted that position in the answer matters almost as much as whether you appear at all. Brands named first in an AI recommendation carry the most weight with readers, so tracking position separately from coverage reveals whether you are improving in quality as well as frequency.

How to build a free prompt tracking spreadsheet

A simple Google Sheets or Excel spreadsheet is all you need. Set it up once and it becomes a running record of your GEO progress.

The key is a structure that makes each session take minutes to fill in, not hours, and that makes trends visible at a glance without any formulas or dashboards.

  1. Create one sheet per engine. Three tabs: ChatGPT, Gemini, Perplexity. Each tab holds the same structure so you can compare across engines easily.
  2. Rows = prompts, columns = sessions. List your 10-15 prompts down the left column. Each time you run a check, add a new date column to the right. This gives you a visual timeline at a glance.
  3. Use a compact encoding per cell. For each prompt-session cell, record: M/P/N/X (Mentioned/Position/No-mention/X for error), a citation URL if present, and a one-letter sentiment (P/N/U for positive, negative, unclear). Example: "M2, example.com/page, P" means mentioned in position 2, that URL was cited, described positively.
  4. Add a summary row at the top. Count the number of M results per session column. That single row, tracked over time, shows your brand coverage trend without any formulas beyond a basic COUNTIF.
  5. Colour-code competitor appearances. In a separate "competitors" column per cell, mark any competitor names. Highlight in yellow when a specific competitor appears. Red patterns in that column show you exactly where you are losing.

Pew Research found 34% of US adults had used ChatGPT by 2025, with product research among the top use cases. The buyers filling those answers are real.

Having even a simple spreadsheet record lets you see whether your content work is moving you in front of them.

How often should you run manual checks?

Every two weeks is the right cadence for most founders doing this manually. Weekly is better but unsustainable alongside a full-time build schedule.

Monthly is the minimum that still gives you a meaningful trend. What you want to avoid is checking once after a content change, seeing no movement, and concluding the change did not work.

AI answers update on different timescales: Perplexity refreshes fastest (days to a week for newly indexed content), ChatGPT's browsing mode is slower and less predictable, and Gemini follows Google's crawl cycle.

The practical rule: make a change (new page, new listing, updated description), wait three weeks, then run a full session. That three-week window gives Perplexity time to index and re-answer, ChatGPT time to browse the updated page, and Gemini time to reflect the Google index update.

Checking sooner than that is how you incorrectly conclude a change did not work when it just has not propagated yet.

For context on what each engine's update cycle looks like, our guide on how to check if AI engines mention your brand covers the differences in retrieval speed across ChatGPT, Gemini, and Perplexity.

What counts as real progress vs random noise?

Real progress is a sustained directional change across at least three consecutive sessions. A single session where you appear in 8 out of 10 prompts after appearing in 5 the session before is not real progress.

It might be. But one session is noise.

Three sessions in a row showing an upward trend, on the same prompts, is signal. The Princeton GEO research found meaningful, reproducible visibility improvements of up to 40% from content optimisations, but those improvements only become visible with consistent measurement over weeks, not days.

Common sources of noise to recognise and filter out:

  • Phrasing variation. A slightly different prompt wording can change which brands appear. If your coverage jumps, check whether you inadvertently asked the prompt differently than last time.
  • Session-to-session engine variability. AI answers fluctuate. Two identical prompts run one hour apart can return different results. Track trends over sessions, not individual answers.
  • New competitor entries. A competitor launching a new product or getting press coverage can briefly dominate answers across your prompts. A sudden drop in your share of voice may be their spike, not your decline.
  • Engine updates. Perplexity, ChatGPT, and Gemini all update their models periodically. A broad shift in results across all your prompts at once often signals a model update, not a change you caused.

When does manual tracking stop being enough?

Manual tracking stops being enough when the time it takes to run sessions consistently starts crowding out the content work that actually moves the numbers. For most founders, that threshold arrives around 15 prompts across three engines, run fortnightly.

That is 45 individual checks per session, plus recording and analysis time. Done properly, it takes 90-120 minutes every two weeks, and it has to be done on a fixed schedule to remain reliable.

The other signal that you have outgrown manual tracking: you start skipping sessions because the friction is too high. A month of missing data breaks your trend line and leaves you with no baseline to compare against when you want to evaluate a specific change.

At that point, the cost of the manual method exceeds the cost of a lightweight tool.

The goal of manual tracking is not to do it forever. It is to build enough of a baseline, over 8-12 weeks, that you understand your own visibility well enough to know what to fix.

Once you have that foundation, our guide on how to track brand visibility in AI search covers what a more systematic measurement approach looks like.

Frequently asked questions

Do I need accounts on ChatGPT, Gemini, and Perplexity to track manually?

Perplexity lets you run queries without an account. ChatGPT and Gemini both work with free accounts. A paid ChatGPT plan gives access to browsing mode, which returns more current live-web results and is worth using if you have it. For manual tracking purposes, free accounts on all three are sufficient to get a meaningful baseline.

Should I track every prompt separately or average them?

Track them separately. Averages hide the most useful information: which specific prompts you are winning, which you are losing, and to whom. A prompt where you consistently appear in position one is different from a prompt where you never appear. Grouping them loses the detail you need to decide which content to work on next.

Can I use a shared spreadsheet with a co-founder to split the tracking work?

Yes, and it helps with consistency. Agree on the exact prompt wording in advance and lock it in the spreadsheet so neither person paraphrases. The biggest risk with shared tracking is prompt drift, where two people ask the "same" question in slightly different ways and get different results, making the data incomparable across sessions.

How do I know if my content changes are actually causing the improvement I see?

Change one thing at a time and give it three to four weeks before evaluating. If you update your G2 listing, publish a new FAQ page, and revise your robots.txt in the same week, and your coverage improves the next session, you will not know which change drove it. Isolating variables is harder in practice than in theory, but changing fewer things per period makes attribution far cleaner.

People also ask

Your customers are asking AI for buying decisions. Is your brand being recommended?

Lead Rescue helps you measure your AI visibility, compare your brand against competitors, and uncover opportunities to grow your brand across ChatGPT, Gemini, and Perplexity.

Get Started For Free