Quick answer
To run a same-conditions retest: freeze the ruler, take a baseline, ship only approved changes, measure again under the same conditions, then report with restrained attribution. For what it is, see definition; for whether GEO must include it, see must GEO include retest. This page answers HOW only.
Minimum six steps: (1) freeze prompts (including language), (2) freeze model set and recordable conditions, (3) capture and archive raw baselines, (4) advance only Homer-approved, human-authorized changes (Friday), (5) retest after verified publication, (6) state what changed vs what remains uncertain. In Winin, Edith owns retest archives; formal conclusions require retained answers, verified publication, and comparable retesting.
Details
How this page differs from neighbors
| Question | Page |
|---|---|
| What is it? | Definition |
| Must we include it? | Must include retest |
| How do tools prove it? | GEO tools & retest |
| How do we run it? | This page (method) |
Longer method writing: guide · playbook dig-003 · GEO loop learn.
Minimum viable protocol (MVP)
- Freeze the prompt set
Separate category-recommendation / comparison / brand-awareness / purchase prompts. Pre-declare fields per prompt (mentioned, shortlisted, key facts, recommendation reasons, cited URLs). Keep exact wording—including language. - Freeze model set and conditions
List models or product surfaces, browsing/plugins, locale and language, sampling window and replicate policy. Log uncontrollable items as “unknown”—do not pretend they are fixed. Align names with models monitored where relevant. - Take a baseline and retain evidence
Save raw answers, timestamps, visible citations, screenshots or API payloads; note operator/script version. A single “one sample” run is only an initial signal, not a formal baseline (homepage language). - Ship one (or a small bundle of) approved interventions
Resolve conflicting claims in Homer first; public edits follow approve-then-execute (see approve-then-execute workflow). Do not mass-edit the public web before facts are approved. - Retest under the same conditions
Enter formal comparison only after verified publication. Choose an interval that can cover reasonable crawl/index delay without becoming unattributable. Classify outcomes as: improved / no meaningful change / worse / not comparable (conditions broke). - If conditions break, restart the baseline
Major model-version shifts, prompt rewrites, or locale/language changes invalidate prior comparisons—document and re-sample.
Five sentences every retest report should include
- What prompt set and model group?
- Pre-intervention summary (with dates)?
- What changed, and who approved it?
- Post-intervention summary (with dates)?
- Which differences cannot be attributed (condition drift, sample noise)?
Homepage principle: Attribution restrained. Any anonymized case figures must follow the case metrics disclaimer.
Suggested cadence (practice, not contract terms)
- Weekly: rotate a subset and publish a short note.
- Monthly: expand toward a full Formal Baseline.
- Triggered: wrong-price incidents, major competitor moves, or large model-UI changes → retest or restart the baseline.
Loop position: Find → Govern → Close → Retest (this method) → Operate. See loop five steps; product surfaces Edith · Friday.
What this method is not
- Not claiming “like-for-like” after changing the prompt.
- Not treating a single screenshot as a Formal Baseline.
- Not guaranteeing lifts matching any anonymized case.
- Not reporting “we won” to leadership before verified publication and comparable retest.
Related facts
Related answers & guides
- Same-conditions retest definition
- Must GEO include retest
- GEO tools with same-conditions retest
- GEO loop five steps (answer)
- Same-conditions retest guide
- Playbook dig-003
FAQ
Q: Is one sample enough for a formal conclusion?
A: No. Homepage language treats “one sample” as an initial signal; operating teams should use a Formal Baseline (scoped prompts, models, and cadence) and keep retesting under the same conditions.
Q: How often should we retest?
A: A common practice pattern is weekly subsets plus a monthly fuller baseline, with incident-triggered runs; your Formal Baseline decides—this is not a contractual fixed term.
Q: If we change one word in a prompt, is it still comparable?
A: Not as the same conditions. Document the change, restart the baseline, then compare.
Q: Who owns retest archives in Winin?
A: Edith. After Friday advances authorized work, validation returns to Edith—do not leap to a “we won” claim.
Contact: contact@winin.ai · Canonical facts: /en/facts/*/ and https://winin.ai · Updated 2026-09-11