RankSurf

METHODOLOGY

How we calculate visibility and confidence

How it works, and where it stops being certain.

Methodology v1.0 · Last reviewed August 2026

TL;DR

  • Every number we show is a range, not a point.
  • We account for the fact that asking the same question repeatedly isn’t the same as asking many different questions.
  • We only call a month-over-month move ‘real’ when it clears a bar sized to your own data volume.
  • We can’t fully separate a real shift in your visibility from an AI model changing its own answers without announcing it. More scans narrow this. They do not eliminate it.

Why a bare percentage would mislead you

A bare number hides how much it doesn’t know.

Political polling got this right decades ago. A poll never reports “52% support” on its own. It reports “52% support, ±3 points, 95% confidence,” because a poll of a thousand people is an estimate of the whole electorate, not a headcount of it. Readers already trust that model. We use the same one, for the same reason: your visibility rate is an estimate too. We ask a limited set of prompts, a limited number of times, whether an AI engine mentions your brand, and from that sample we infer the rate across every way a real person might ask. A bare percentage on a dashboard implicitly claims more precision than that sample can support. We don’t print one without saying how sure we are.

This cuts in the direction you’d least expect, too. If we check a prompt 28 times and your brand never once shows up, we don’t print a flat 0%. Twenty-eight checks finding nothing is meaningfully different from an infinite number of checks finding nothing. The honest statement is that the true rate is probably below roughly 12%, not that it’s impossible. Zero observed mentions is not the same claim as zero probability, and our numbers say so.

That case is not a corner case. Across 126 buyer questions we track for real brands, covering roughly 3,000 checks in a single month, the brand never appeared even once in 58% of them. Absent from the answer, not buried in it. If that turns out to describe you, you are in the majority, and the number that matters is not today’s percentage but whether it moves. Which is the rest of this page.

How we turn one number into a range

More evidence narrows the range, less evidence widens it, and that is the whole mechanism.

Picture guessing whether a coin is fair. Flip it 5 times and get 3 heads, 2 tails. That tells you almost nothing. The result fits a wide range of true probabilities, from a coin that’s clearly biased to one that isn’t. Flip it 500 times and land close to 250/250, and the plausible range around “fair” gets tight fast. Nothing about the coin changed between those two experiments. Only how much evidence you had.

Each time we check whether a prompt gets your brand mentioned by an engine, that’s one flip. The more independent checks feed an estimate, whether from more prompts tracked or more repeat scans over the window, the tighter the plausible range around the true rate gets. Fewer prompts or a lighter scan schedule means a wider range, for the same underlying reason a 5-flip coin test tells you less than a 500-flip one. A wide range at low evidence is the math refusing to pretend it knows more than it does.

Worked example

Take one real production account we monitor. Over a 28-day window it tracked 13 buyer-question prompts, checked 187 times between them, and the brand showed up in about 50% of the AI answers. The honest range around that 50% runs from 34% to 67%. That’s not the model breaking. Thirteen prompts, even checked repeatedly, is a modest amount of evidence about how an engine behaves across every way a buyer might phrase a question.

One thing worth knowing: your brandedprompts stay out of this number, and out of every report we generate for you. Questions like “what is [your brand]” almost always mention you. Counting that as visibility would be circular, and folding them in would inflate the headline dramatically on most accounts. You can switch to a branded view inside the app if you want to see it, and we will tell you there what it is worth. It stays out of anything you hand a client. Track more prompts, or scan more often, and the range tightens. The true rate has not moved. There is simply more evidence standing behind the estimate.

Why we don’t treat every repeat check as independent

Asking the same question five times isn’t the same as asking five different questions.

This is the part of the method we’d most want a skeptical reader to check, because it’s the part that makes the honest number look worse than the convenient one. If an AI engine’s answer to “best CRM for startups” mentions you today, it’s very likely to keep mentioning you tomorrow, because the underlying answer text isn’t reshuffling itself between checks the way genuinely independent samples would. Counting 20 repeat checks of that one prompt as 20 independent data points is like calling one friend and asking the same question five times, then treating that as five opinions. You collected one opinion, repeated five times, dressed up as more evidence than it is.

The naive fix is to pool every single check across every prompt into one giant tally and run the standard math on that. It understates the true error for exactly this reason, and it does so by a lot, not a little. On one real production account we monitor, covering 13 buyer-question prompts and 187 completed checks over a 28-day window, the honest interval, computed by treating whole prompts as the unit of repetition, runs from 34% to 67%: about 34 percentage points wide. The naive calculation, pooling every check as if it were an independent observation, would have printed 45% to 59%, an interval only 14 pointswide, on the exact same underlying data. That’s well over twice as precise as the evidence supports, and we see the same pattern on every account we check. That gap is not rounding. It separates a range you could defend to a client from one that looks impressively tight because it is measuring the wrong thing.

Every visibility number on RankSurf treats the whole prompt as the unit of repetition, so checks of the same question are never mistaken for independent evidence, even when that means showing you a wider, less flattering range than the naive method would. Within a prompt we do still account for the ordinary luck of individual checks; what we never do is let those checks pose as separate questions.

How we decide a change is real

We compare each prompt against itself, then ask how often pure chance could produce a swing that large.

For every prompt tracked in both periods, we take the difference between this month’s rate and last month’s. That gives us a set of per-prompt movements, some up, some down. Then we ask a blunt question: if none of these movements were real, and the direction of each one were as arbitrary as a coin flip, how often would shuffling those directions at random produce an overall shift as large as the one we actually saw? If the answer is “rarely, less than one time in twenty,” we call the move real. Otherwise we say it’s within noise.

It’s worth being specific about what we don’tdo, because the obvious shortcut is wrong in a way that flatters us. You might expect us to check whether this month’s range and last month’s overlap, and call it real when they don’t. We tested that. It declares noise “real” far more often than its own 95% label promises, and it gets worse the more prompts an account tracks, which is exactly backwards. A method that becomes less trustworthy as a customer invests more in it is not one we were willing to ship, so the ranges above are for reading; the verdict comes from the shuffle test.

We run this on ourselves too. Our own RankSurf account had 9 buyer-question prompts tracked in both windows. At that count, our simulations show this design can only reliably detect a swing larger than 10.5 percentage points, so that is the bar we hold our own numbers to. Our visibility did nudge upward over that period, by a fraction of a point. Nowhere near the bar. The honest verdict on our own dashboard therefore read “within measurement noise,” not “real.” We could have written the code to call that a win. It would rather tell us the truth than hand us a nicer-sounding story about our own product, and the same is true of the number it shows you.

That detection floor, the smallest change a given setup can reliably tell apart from noise, sharpens with more prompts and denser scanning. We print it beside every “within noise” verdict, because that is where it changes the meaning of what you’re reading: “we couldn’t confirm this” is a very different statement depending on whether the bar was 4 points or 12. When a change isconfirmed we show its plausible range instead, which tells you the same thing more directly. We’re not printing either one to talk anyone into tracking more prompts. It’s simply the true state of the evidence behind your number, and that number happens to move with your setup.

One more choice worth stating outright, because it’s deliberate rather than an oversight: we compare full 28-day windows against each other, not single-day snapshots. That means whatever we call “noise” includes ordinary day-to-day drift, including the possibility that an AI engine changed its own answers between one scan and the next without announcing it. We build that in on purpose. The question this comparison answers is “is this month different from last month,” and for that question the honest noise floor has to include the everyday wobble your engines already show on their own, not just the sampling error you’d get from one clean-room snapshot. A design that stripped that wobble out would call real change more often than it should. This one won’t.

What this can’t tell you

A real change is real evidence something shifted. It is not proof of what caused it.

We can’t fully separate “your visibility changed” from “the AI model changed.” AI engines update their own training, tuning, and retrieval on a schedule that’s invisible to everyone outside the company that runs them. When your rate moves in a way we call “real,” that tells you something shifted in what these engines say about you. It does not tell you whether to credit your own marketing, a competitor’s, or an unrelated model update on the vendor’s side. More frequent scanning narrows this ambiguity over time. It doesn’t eliminate it, because the model itself stays a black box to us, same as it is to you.

A confirmed real change is statistically real, not proof of what caused it. “Real” is a statement about the data: the movement is unlikely to be noise. It is not a story about why it happened. Attributing a real change to a specific campaign, a competitor’s move, or a particular piece of content takes the same judgment any analyst would apply to correlated evidence. We don’t manufacture a causal claim the measurement can’t support.

Small and low-cadence accounts will have wide ranges for a long time, by design. If you’re tracking a handful of prompts on a light scan schedule, your honest range will stay wide, sometimes uncomfortably so, for weeks or months, because that little evidence genuinely can’t support a narrower one. We will not shrink it artificially to look more finished than the evidence allows. The fix, when you want one, is more prompts or a denser scan schedule. The formula is not the thing holding you back.

The assumptions behind all of this are periodically re-validated against our own production data, not left to go stale in a spec document. The version stamp at the top of this page reflects that review, and we’ll update it whenever the method changes.

We’d rather hand you a range you can defend in a client meeting than a clean-looking number that overstates what we know. If something on this page doesn’t add up, or you want to see this math applied to your own account, email us at hello@ranksurf.com and we’re glad to walk through it.