You measure AI visibility by treating it as a trend, not a snapshot. Run the same fixed set of prompts across the same tools at least three times, log the answers and their sources every run, and score how often you appear rather than whether you appeared once. A single answer proves nothing, because AI gives a different response each time you ask. A pattern across several runs is the only number worth acting on, and screenshots are what turn that number into evidence you can compare later.
This is the discipline most DIY tests skip. Someone asks ChatGPT once, sees their name, and declares victory, or asks once, doesn’t see it, and panics. Both readings are noise. If you’ve already run a first check using how to check if AI recommends your business, this is the next layer: turning scattered impressions into a measurement you can trust.
How do I measure AI visibility when the answers keep changing?
Hold everything constant except time. Use the same prompts, the same tools, the same clean-session conditions, and run them in batches. Then count your appearance rate across the batch instead of reacting to any one answer. Measuring the same thing repeatedly is what lets you separate a real change from the model’s normal wobble.
The mindset shift is the whole game. AI visibility behaves less like a Google ranking you can look up and more like a weather pattern you sample. You’re not asking “am I visible right now.” You’re asking “across ten tries, how often does AI put me in the answer, and is that share going up or down over the months.” Once you frame it that way, the changing answers stop being a problem and become the thing you’re measuring.
Why does AI give a different answer every time?
Because language models are built to vary. They don’t retrieve one fixed record; they generate a fresh response each time, with a degree of randomness by design. On top of that, tools like ChatGPT and Perplexity pull live web results that themselves change, and signed-in sessions personalize based on your history. Three sources of variation, stacked.
That’s why one measurement is meaningless and why cleaning up your test conditions matters so much. Run signed out or incognito to strip out personalization. Ask each prompt in its own fresh conversation so answers don’t influence each other. You can’t remove the built-in randomness, and you shouldn’t try. You measure across it.
What should I actually measure?
Track four things per prompt, run after run. Presence is the headline; the rest add texture and tell you what to fix. Keep it in a spreadsheet so the numbers accumulate into a trend instead of evaporating after each session.
| Metric | What it tells you | How to score it |
|---|---|---|
| Presence rate | How reliably AI names you at all | Times you appeared ÷ total runs (e.g. 3/10 = 30%) |
| Average position | Whether you’re a first pick or an afterthought | Note your spot in the list each time, then average |
| Share of voice | How you stack up against named competitors | Your mentions ÷ all business mentions across runs |
| Source cited | Which pages AI reads about you, if any | Log the URLs Perplexity and ChatGPT show as sources |
Presence rate is the number to watch month over month. If it climbs from 20% to 60% after you fix a set of signals, that’s a result you can defend. If it doesn’t move, you’ve learned something too: the change you made wasn’t the lever.
How many runs, and how often?
Run each prompt at least three times per measurement cycle, and repeat the whole cycle monthly. Three runs is the floor for telling signal from noise; more is better if you have the patience. Monthly cadence matches how slowly the underlying signals and re-crawls actually move, so you’re not mistaking day-to-day randomness for progress.
A practical rhythm looks like this:
- Set a fixed prompt list and don’t change it, or you lose comparability.
- Run the full set three times in a measurement cycle, across ChatGPT, Gemini, Perplexity, and Claude.
- Log every answer and source, and screenshot each one with the date visible.
- Calculate your presence rate and note position and competitors.
- Wait a month, change one thing at a time in between, and run the identical set again.
- Compare cycles, not sessions. The line between months is your real scoreboard.
Changing one thing at a time is what makes the measurement diagnostic. If you unblock the crawlers, rewrite three pages, and fix your directories all at once, a later bump tells you something worked but not what. Space the changes out and the trend line tells you which move mattered.
Turn it into a baseline you can compare
Your first full cycle is your baseline: the honest before-picture you’ll measure every future change against. Date it, save the screenshots, and record the exact conditions you used. Without that anchor, “I think it’s better now” is all you’ll ever have, and that’s not a number you can make decisions with.
Screenshots matter more than they seem. You cannot reproduce today’s exact answer next month, so the image is your proof. When a client or a partner asks whether the work paid off, a dated before-and-after of the same prompt across three runs is far more convincing than any dashboard claim. Evidence beats assertion, especially in a medium this fluid.
Where the DIY method hits its limit
Hand it to a pro when you need this measured rigorously, repeatedly, across all four tools, with the causes explained. Running three cycles a month by hand across four tools is a real time cost, and the manual test still can’t see what the model reads inside your site, or reliably reproduce identical conditions. For tracking your own trend casually, the method above is genuinely enough.
For a benchmark you can build on, our AI visibility audit runs the same set of questions across ChatGPT, Gemini, Perplexity, and Claude, logs the answers and sources, and turns them into a score out of 100 across 39 criteria, plus a prioritized fix list. That score becomes your baseline, and the Report + Implementation Plan package includes a re-measurement 30 days later so you can see the trend without running every cycle yourself. It starts at PLN 499 net (around $125), with a refund if we don’t find at least five things worth fixing. Once you know you have a gap, why ChatGPT doesn’t recommend your business and, for local trades, AI visibility for small and local businesses cover what to change. Prefer to talk first? Book a free 20-minute consultation or start with our free 3-minute marketing audit.
FAQ
How many times should I ask before I trust the result? At least three runs of the same prompt per measurement cycle, and more if you can. Once is meaningless because of the built-in randomness in AI answers. Three lets you calculate a rough presence rate; five to ten makes it steadier. The exact number matters less than being consistent: use the same count every cycle so your months are comparable.
Is there a single AI visibility score I should track? Presence rate is the most useful single number: how often, out of all your runs, AI names your business. It’s simple, it maps directly to being recommended, and it moves when your signals improve. Position, share of voice, and cited sources add helpful detail, but if you track only one thing over time, track presence rate as a percentage.
Can I use a tool to measure this automatically? Some paid trackers monitor AI mentions for you, which saves the manual runs. They’re useful for the counting, but a tool measures; it doesn’t explain why you’re absent or which fix comes first. If you use one, pair the numbers with a human read of your site. We cover that trade-off, and where an audit fits, on the AI visibility audit page.
How long before a change shows up in my measurements? Plan on weeks to a couple of months, and don’t credit a change too early. Technical fixes like unblocking crawlers can register once AI tools re-crawl you; trust signals like reviews accrue slowly. That’s why the cadence is monthly and why you change one thing at a time: it lets the trend line, not a lucky single answer, tell you whether the work landed.