RankSurf

METHODOLOGY

How we calculate visibility and confidence

How it works, and where it stops being certain.

Methodology v1.2 · Last reviewed September 2026

TL;DR

  • Every number we show is a range, not a point.
  • We account for the fact that asking the same question repeatedly isn’t the same as asking many different questions.
  • We run two different tests against the same evidence and let the stronger one decide, only calling a month-over-month move real once it clears a bar sized to your own data volume.
  • We can’t fully separate a real shift in your visibility from an AI model changing its own answers without announcing it. More scans narrow this. They do not eliminate it.
  • Our simulated error rate hasn’t yet been checked against enough of our own real history to call it proven, and we don’t yet run the negative-control prompts the fuller design calls for.
  • For Google Search traffic we run a different design entirely — a controlled interrupted time series comparing your page group against frozen comparison sections — and on our own data it withholds a verdict far more often than it renders one. Its error rate is checked only against simulated noise so far, not yet our own history.

Why a bare percentage would mislead you

A bare number hides how much it doesn’t know.

Political polling got this right decades ago. A poll never reports “52% support” on its own. It reports “52% support, ±3 points, 95% confidence,” because a poll of a thousand people is an estimate of the whole electorate, not a headcount of it. Readers already trust that model. We use the same one, for the same reason: your visibility rate is an estimate too. We ask a limited set of prompts, a limited number of times, whether an AI engine mentions your brand, and from that sample we infer the rate across every way a real person might ask. A bare percentage on a dashboard implicitly claims more precision than that sample can support. We don’t print one without saying how sure we are.

This cuts in the direction you’d least expect, too. If we check a prompt 28 times and your brand never once shows up, we don’t print a flat 0%. Twenty-eight checks finding nothing is meaningfully different from an infinite number of checks finding nothing. The honest statement is that the true rate is probably below roughly 12%, not that it’s impossible. Zero observed mentions is not the same claim as zero probability, and our numbers say so.

That case is not a corner case. Across 126 buyer questions we track for real brands, covering roughly 3,000 checks in a single month, the brand never appeared even once in 58% of them. Absent from the answer, not buried in it. If that turns out to describe you, you are in the majority, and the number that matters is not today’s percentage but whether it moves. Which is the rest of this page.

How we turn one number into a range

More evidence narrows the range, less evidence widens it, and that is the whole mechanism.

Picture guessing whether a coin is fair. Flip it 5 times and get 3 heads, 2 tails. That tells you almost nothing. The result fits a wide range of true probabilities, from a coin that’s clearly biased to one that isn’t. Flip it 500 times and land close to 250/250, and the plausible range around “fair” gets tight fast. Nothing about the coin changed between those two experiments. Only how much evidence you had.

Each time we check whether a prompt gets your brand mentioned by an engine, that’s one flip. The more independent checks feed an estimate, whether from more prompts tracked or more repeat scans over the window, the tighter the plausible range around the true rate gets. Fewer prompts or a lighter scan schedule means a wider range, for the same underlying reason a 5-flip coin test tells you less than a 500-flip one. A wide range at low evidence is the math refusing to pretend it knows more than it does.

Worked example

Take one real production account we monitor. Over a 28-day window it tracked 13 buyer-question prompts, checked 187 times between them, and the brand showed up in about 50% of the AI answers. The honest range around that 50% runs from 34% to 67%. That’s not the model breaking. Thirteen prompts, even checked repeatedly, is a modest amount of evidence about how an engine behaves across every way a buyer might phrase a question.

One thing worth knowing: your brandedprompts stay out of this number, and out of every report we generate for you. Questions like “what is [your brand]” almost always mention you. Counting that as visibility would be circular, and folding them in would inflate the headline dramatically on most accounts. You can switch to a branded view inside the app if you want to see it, and we will tell you there what it is worth. It stays out of anything you hand a client. Track more prompts, or scan more often, and the range tightens. The true rate has not moved. There is simply more evidence standing behind the estimate.

Why we don’t treat every repeat check as independent

Asking the same question five times isn’t the same as asking five different questions.

This is the part of the method we’d most want a skeptical reader to check, because it’s the part that makes the honest number look worse than the convenient one. If an AI engine’s answer to “best CRM for startups” mentions you today, it’s very likely to keep mentioning you tomorrow, because the underlying answer text isn’t reshuffling itself between checks the way genuinely independent samples would. Counting 20 repeat checks of that one prompt as 20 independent data points is like calling one friend and asking the same question five times, then treating that as five opinions. You collected one opinion, repeated five times, dressed up as more evidence than it is.

The naive fix is to pool every single check across every prompt into one giant tally and run the standard math on that. It understates the true error for exactly this reason, and it does so by a lot, not a little. On one real production account we monitor, covering 13 buyer-question prompts and 187 completed checks over a 28-day window, the honest interval, computed by treating whole prompts as the unit of repetition, runs from 34% to 67%: about 34 percentage points wide. The naive calculation, pooling every check as if it were an independent observation, would have printed 45% to 59%, an interval only 14 pointswide, on the exact same underlying data. That’s well over twice as precise as the evidence supports, and we see the same pattern on every account we check. That gap is not rounding. It separates a range you could defend to a client from one that looks impressively tight because it is measuring the wrong thing.

Every visibility number on RankSurf treats the whole prompt as the unit of repetition, so checks of the same question are never mistaken for independent evidence, even when that means showing you a wider, less flattering range than the naive method would. Within a prompt we do still account for the ordinary luck of individual checks; what we never do is let those checks pose as separate questions.

How we decide a change is real

We run two tests against the same evidence and let the one built to catch small, real moves make the call.

There are two ways to ask whether a month-over-month change is real, and they disagree about how much evidence they need to see it. The primary one, the one that actually decides your verdict, treats every single AI-answer check as its own data point, not just each prompt’s average, so it can pick up a real move from a fairly small set of prompts as long as each one was checked enough times. The second method only ever looks at the handful of per-prompt averages themselves and asks how often shuffling their direction at random, as if each were a coin flip, would produce a shift this large by chance. That second method has a hard floor built into its own arithmetic: with five or fewer prompts whose rate actually moved between periods, it cannot reach statistical significance at any effect size, however large. Below that floor, staying quiet does not mean nothing happened. It means the test was never able to see.

We run both, on the same underlying numbers, and we show both. The first is the primary test and the one that sets your verdict. The second runs alongside it as a cross-check, its own result and p-value shown next to the primary one, but it never gets a veto. Requiring the weaker test to agree before the stronger one is allowed to speak would put the floor described above back in place for every account, which defeats the reason we run two tests instead of one.

Before either test runs, we check there is enough raw material to ask the question honestly: at least 8 prompts tracked in both periods, and a running median of at least 4 completed checks per prompt, in each period. Short of either bar we don’t guess and we don’t call it noise either — we say plainly that there isn’t enough evidence yet, and name exactly what would close the gap: more prompts, more checks, or both.

When the primary test does call a move real, we also show the plausible range for how big that move was — built the same way as every other range on this page, by resampling whole prompts, never individual checks, for the reason the section above lays out.

Both tests come with their own published sensitivity table, generated by simulation against our real prompt-mix data, not a closed-form formula — an earlier formula we tried badly understated a real cohort’s precision, in the wrong direction to be wrong in. As one example straight from that table: a cohort of 11 prompts checked about 12 times each per period can reliably detect a swing above roughly 11.5 percentage points using the primary test; the same 11 prompts checked only 4 times each per period needs a swing above roughly 24.5 points before the design can tell it apart from noise. More prompts and denser scanning both push that floor down, and we print the actual floor for your own setup beside every verdict rather than one number for every account.

The same pair of tests is what decides whether a specific piece of delivered work moved the number, not only the month-over-month figure. That measurement window starts only once the work is actually live on your site, never when it’s merely planned, and it is never extended once it’s running: choosing a longer window after you can already see how things are going would let us pick the answer we wanted. If something else changes in the same area of the site during that window, or a Google update lands in the middle of it, we withhold the causal claim and keep showing the estimate with a note about why, rather than credit your work for someone else’s change. Where we can confirm the exact page a delivered fix touched, we only withhold the claim if the new change hit that same page; for checks that don’t involve fetching a live page, and so don’t give us a specific page to match against, we fall back to the general area of the site the work targeted, which rules a few more windows out than strictly necessary. We chose the direction that under-claims a real win over the one that overclaims a fake one.

What this can’t tell you

A real change is real evidence something shifted. It is not proof of what caused it, and two of the limits below are about our own measurement, not just about what it can tell you.

We can’t fully separate “your visibility changed” from “the AI model changed.” AI engines update their own training, tuning, and retrieval on a schedule that’s invisible to everyone outside the company that runs them. When your rate moves in a way we call “real,” that tells you something shifted in what these engines say about you. It does not tell you whether to credit your own marketing, a competitor’s, or an unrelated model update on the vendor’s side. More frequent scanning narrows this ambiguity over time. It doesn’t eliminate it, because the model itself stays a black box to us, same as it is to you. For the Google side specifically, we watch Google’s own published incident feed for acknowledged issues and rollouts, not a hand-curated calendar of every ranking update, because no complete one is public. A quieter change Google never announces will not show up here.

A statistically real change is not the same as a caused one. “Real” is a statement about the data: the movement is unlikely to be the kind of noise our tests are built to catch. It is not a story about why it happened. Attributing a real change to a specific campaign, a competitor’s move, or a particular piece of content takes the same judgment any analyst would apply to correlated evidence. We don’t manufacture a causal claim the measurement can’t support — and when the claim is attached to one specific piece of delivered work rather than a month-over-month figure, we withhold it entirely, rather than soften it, whenever something else changed in the same window that could explain the move instead.

Our own false-positive rate is simulated, not yet checked against our own history. We know how often each test would wrongly call a real change under a model of what “nothing happening” looks like, but that model is a simulation built from an idealized picture of how independent your checks are. The real test is running the same design against long stretches of our own genuinely intervention-free history and seeing how often it cries wolf there instead. We’ve built that check, but we don’t yet have enough usable history to run it with any confidence: roughly six clean, non-overlapping stretches across our current accounts, against the twenty or more a false-positive rate needs to tell a 5% rate apart from a 15% one. Until we have enough, treat the “1 in 20” behind every “real” verdict as our best current estimate of the design’s honesty, not as a number already proven against reality. We’ll publish the result, and move the version stamp above, the day we have it.

We don’t yet run negative-control prompts. The fuller version of this design tracks a handful of prompts your work couldn’t plausibly affect, checked in the same runs as the ones that matter, specifically to catch a move that looks real but was actually the AI model shifting under everyone at once rather than anything you did. That safeguard is specified and not yet built. Until it is, a verdict on a small cohort carries a bit more of that risk than the rest of this page implies, which is one more reason, on top of the ones above, that more prompts and more frequent checks make a verdict more trustworthy in the meantime, not just faster.

Small and low-cadence accounts will have wide ranges for a long time, by design. If you’re tracking a handful of prompts on a light scan schedule, your honest range will stay wide, sometimes uncomfortably so, for weeks or months, because that little evidence genuinely can’t support a narrower one. We will not shrink it artificially to look more finished than the evidence allows. The fix, when you want one, is more prompts or a denser scan schedule. The formula is not the thing holding you back.

The two limits above aren’t a footnote — they’re why this page says “simulated” rather than “proven,” and they’re also a to-do list, not a permanent state. We’ll update this page, and the version stamp at the top, the moment either one closes.

How we measure whether a change moved Google Search traffic

This is a different design from everything above — daily click counts, not a mention rate — and it withholds a verdict more often than it renders one, including on our own account.

For Google Search Console outcomes we don’t run either of the two tests described above. We fit a controlled interrupted time series (ITS): a quasi-Poisson regression, estimated by iteratively reweighted least squares, over daily click counts for the page group you changed and a small set of frozen comparison sections chosen before the measurement window opened. The regression includes day-of-week terms — traffic to any site follows a weekly rhythm, and ignoring that would read an ordinary Tuesday dip as movement — and each group’s own pre-existing trend, so what’s left over, the interaction between “this is the changed group” and “this is after the change,” is the one number we report: the log rate-ratio in daily clicks the comparison design assigns specifically to your change, net of whatever moved the comparison sections too.

Organic click counts are noisier day to day than an idealized count model expects. Rather than assume that noise away, we measure how much extra dispersion the data itself shows and widen every interval by exactly that amount, so a volatile page group gets an honestly wider range instead of borrowing false precision from a calmer one.

The comparison sections are chosen once, at the moment the measurement clock starts, purely on how closely their own traffic and page count matched the group you changed before the change happened — never on a backlink count, a topic-similarity score, or any other stand-in for “this seems related.” We keep watching them for the rest of the window: if the comparison sections move materially on their own — a site-wide change, a template update, anything not specific to the group you actually touched — we withhold the verdict rather than credit your work with a shift the whole site experienced. That check judges the comparison sections as a pooled group, using only how much they naturally bounced around before the change to decide how big a shift counts as material, never the period after, since that is the stretch that might already have moved.

A verdict is withheld, and the underlying estimate still shown beneath it, whenever any of the following is true:

  • Not enough traffic yet. Roughly 1,000 clicks a day to the group, or roughly 1,000 pages in it — cleared by either one, not both. Below that floor the arithmetic itself can’t separate a real shift from ordinary daily noise, however the code is tuned. Most sites don’t clear it, ours included, so this is the majority outcome here, not an edge case.
  • Not enough history before the change. Fewer than 28 usable days in the pre-period.
  • Too few comparison sections. Fewer than two eligible sections could be frozen as controls when the window opened.
  • The comparison sections moved. They shifted materially over the same window, so we can’t tell your change apart from whatever moved the rest of the site.
  • We had nothing to check an update against. The Google core/spam-update calendar has no coverage for this window (see below) — that is never read as clean.
  • The regression itself failed to fit. A singular design, a fit that never converged, or a treated group with zero clicks throughout. We report that we couldn’t compute a number, never one that merely looks plausible.

On the Google-update check specifically: we keep a calendar of past core and spam-update rollouts and withhold whenever one overlaps your measurement window. That calendar currently only covers rollouts we could confirm through mid-2025 — we deliberately did not extend it further without a live source to check against, rather than guess at dates for anything more recent. A window landing after that point is not read as “clean”; it is read as “we had nothing to check it against,” and withheld on that basis until the calendar is extended.

The false-positive rate behind this design has not yet been measured against our own real history. We ran the estimator 1,000 times against simulated null data with no injected effect — synthetic noise, not real clicks — and it called a false positive 5.30% of the time, close to the 5% a correctly-sized test should produce at that significance level. That confirms the arithmetic is sized correctly against synthetic noise. It is not the same claim as checking it against our own real search history, where clicks to the same pages likely share cache and ranking state across days in ways synthetic noise doesn’t reproduce, so the honest real-world rate could differ from 5% in either direction until we run that check. Doing so needs roughly twenty clean, intervention-free real stretches of history across paying accounts; we currently have far fewer than that. Until we do, treat every search-side verdict this section describes as sized against simulated noise, not yet proven against our own history — the same caveat this page already makes for the AI-answer verdict above, and just as unresolved here.

We’d rather hand you a range you can defend in a client meeting than a clean-looking number that overstates what we know. If something on this page doesn’t add up, or you want to see this math applied to your own account, email us at hello@ranksurf.com and we’re glad to walk through it.