The short answer
One run can reveal an example; it cannot establish stability. For an exploratory check, Really Good GEO uses repeated runs to see whether the same result survives immediate repetition. For a client claim, benchmark, or causal conclusion, choose the sample size from a stated precision target and report uncertainty. “Three runs” or “five runs” is a workflow choice, not a universal statistical guarantee.
The number of prompts and the number of repetitions answer different problems. More prompts broaden coverage across buyer questions. More repetitions estimate how stable the result is for the same question and conditions.
Match the test to the decision
| Use case | What the test can support | Run approach |
|---|---|---|
| Exploratory page check | Find candidate gaps and examples | Repeat each important prompt enough to detect obvious instability; label the result directional. |
| Operational monitoring | Track whether an observed state is recurring | Use the same frozen prompts and conditions on a consistent schedule; retain every run. |
| Before-and-after page test | Compare observations around one controlled change | Collect repeated baseline and post-change runs with matched conditions and a no-change comparison when feasible. |
| Published benchmark or performance claim | Estimate a rate or difference with stated uncertainty | Define the target precision, sampling plan, exclusions, and analysis before collecting results. |
A practical repeated-run protocol
- Write the decision the result will inform.
- Freeze the exact prompts and classify their intent.
- Name the product, visible mode, market, account state, and test window.
- Run the same prompt repeatedly without editing it between attempts.
- Record mentions, visible citations, cited URLs, competitors, and answer accuracy for every run.
- Separate immediate repetitions from later monitoring windows.
- Report the count of observed states, not only a blended score.
The AI Citation Test Log provides the fields needed to preserve those observations.
Use a stopping rule, not a convenient result
Decide the run count or statistical stopping rule before seeing whether the brand wins. Stopping after a favorable answer—or adding runs only after an unfavorable one—changes the meaning of the result.
For higher-stakes reporting, estimate uncertainty instead of relying on a fixed folk rule. The paper Quantifying Uncertainty in AI Visibility treats citation metrics as sample estimators and shows why repeated sampling and confidence intervals matter. Its sample-size findings depend on its platforms, topics, and design; they should not be copied as a universal prescription.
What every result should report
- Exact prompt set and intent classification
- AI product, visible mode, market, and dates
- Runs per prompt and total eligible responses
- Mentions, visible citations, and cited URLs by run
- Exclusions, failures, and missing observations
- Baseline, intervention, and comparison conditions
- Observed variation and any uncertainty interval
- Limits on generalizing beyond the measured test
A run count is not a quality badge. A transparent five-run log can be more useful than a large undocumented batch, while a five-run log may still be inadequate for a benchmark or a narrow difference.
For the language to use around individual results, see why one ChatGPT result is not a ranking.
Freeze the page before testing the outcome.
Audit the source page, save the version, and then run the predetermined question set.
Audit a page free