Method note—not a performance promise. Retest proves change under comparable observation, not permanent model recommendation.
Definition (align language first)
Same-conditions retest: Under as-fixed-as-possible question sets, model/engine groups, prompt and sampling settings, locale/language, and time-window recording, sample a round (or more) before and after an intervention, then compare pre-declared fields such as mention / shortlist / accuracy / citation.
Product loop wording: Find → Govern → Close → Retest → Operate (/en/facts/loop-find-govern-close-retest-operate/). Definition answer: same-conditions-retest-definition.
Why it is needed
Generative answers drift. When communities ask how to reliably measure ChatGPT visibility, the hard part is often not “is there a score,” but whether the same ruler works next time (src-010). Without a retest protocol, optimization weekly reports become screenshot contests.
Minimum viable protocol (MVP)
- Freeze the question set — separate category / comparison / brand-awareness / purchase prompts; declare expected observation fields per question.
- Freeze model group and conditions — engine/model names, API or UI path, login, browsing/plugins; record locale, language, temperature where controllable; mark unknowns instead of pretending they are fixed.
- Capture baseline and evidence — raw text, timestamps, visible citations, screenshots or API payloads; note sampler/script version.
- Apply one (or a small bundle of) approved interventions — fact governance first (/en/facts/homer/); public changes approve then execute (approve-then-execute-geo-workflow, /en/facts/friday/).
- Retest under the same conditions — interval long enough for crawl/index lag expectations, not so long attribution collapses; report improve / no meaningful change / worse / incomparable (conditions broke).
- If conditions break, reopen baseline — major model version shifts, rewritten prompts, locale changes void old comparisons.
Five sentences every report must include
- What were the question set and model group?
- Pre-intervention observation summary (with dates)?
- What intervention, who approved it?
- Post-intervention observation summary (with dates)?
- Which differences cannot be attributed (condition changes, sample noise)?
Forbidden: treating a one-off mention as a Formal Baseline public conclusion (/en/facts/case-metrics-disclaimer/).
How to check tools for real retest support
Ask vendors or your build checklist:
- Lock the same prompt set and re-run in one click?
- Archive raw text side-by-side across runs?
- Expose a “condition fingerprint” (model, time, parameters)?
Related: geo-tools-same-conditions-retest. Monitoring landscape: src-002, methods src-014—neither replaces your protocol design.
Connection to other loop steps
| Step | Retest role |
|---|---|
| Find (Edith) | Baseline and gap location |
| Govern (Homer) | Ensure comparisons use approved facts, not random copy |
| Close (Friday) | Execute only approved changes to reduce attribution noise |
| Retest | This playbook |
| Operate | Fold validated question sets into weekly rhythm |
Overview: geo-loop-five-steps.