Citation testing

How many runs does an AI citation test need?

There is no universal minimum. The right run count depends on whether you are exploring, monitoring, comparing, or publishing a claim.

The short answer

One run can reveal an example; it cannot establish stability. For an exploratory check, Really Good GEO uses repeated runs to see whether the same result survives immediate repetition. For a client claim, benchmark, or causal conclusion, choose the sample size from a stated precision target and report uncertainty. “Three runs” or “five runs” is a workflow choice, not a universal statistical guarantee.

The number of prompts and the number of repetitions answer different problems. More prompts broaden coverage across buyer questions. More repetitions estimate how stable the result is for the same question and conditions.

Match the test to the decision

Use caseWhat the test can supportRun approach
Exploratory page checkFind candidate gaps and examplesRepeat each important prompt enough to detect obvious instability; label the result directional.
Operational monitoringTrack whether an observed state is recurringUse the same frozen prompts and conditions on a consistent schedule; retain every run.
Before-and-after page testCompare observations around one controlled changeCollect repeated baseline and post-change runs with matched conditions and a no-change comparison when feasible.
Published benchmark or performance claimEstimate a rate or difference with stated uncertaintyDefine the target precision, sampling plan, exclusions, and analysis before collecting results.

A practical repeated-run protocol

  1. Write the decision the result will inform.
  2. Freeze the exact prompts and classify their intent.
  3. Name the product, visible mode, market, account state, and test window.
  4. Run the same prompt repeatedly without editing it between attempts.
  5. Record mentions, visible citations, cited URLs, competitors, and answer accuracy for every run.
  6. Separate immediate repetitions from later monitoring windows.
  7. Report the count of observed states, not only a blended score.

The AI Citation Test Log provides the fields needed to preserve those observations.

Use a stopping rule, not a convenient result

Decide the run count or statistical stopping rule before seeing whether the brand wins. Stopping after a favorable answer—or adding runs only after an unfavorable one—changes the meaning of the result.

Directional workflow: for a low-stakes exploratory audit, define a small repeated-run set in advance and describe it as directional. If those runs disagree, the correct conclusion is “unstable in this test,” not a selectively chosen win or loss.

For higher-stakes reporting, estimate uncertainty instead of relying on a fixed folk rule. The paper Quantifying Uncertainty in AI Visibility treats citation metrics as sample estimators and shows why repeated sampling and confidence intervals matter. Its sample-size findings depend on its platforms, topics, and design; they should not be copied as a universal prescription.

What every result should report

  • Exact prompt set and intent classification
  • AI product, visible mode, market, and dates
  • Runs per prompt and total eligible responses
  • Mentions, visible citations, and cited URLs by run
  • Exclusions, failures, and missing observations
  • Baseline, intervention, and comparison conditions
  • Observed variation and any uncertainty interval
  • Limits on generalizing beyond the measured test

A run count is not a quality badge. A transparent five-run log can be more useful than a large undocumented batch, while a five-run log may still be inadequate for a benchmark or a narrow difference.

For the language to use around individual results, see why one ChatGPT result is not a ranking.

Freeze the page before testing the outcome.

Audit the source page, save the version, and then run the predetermined question set.

Audit a page free