← Research & Method

Answers

How to run a same-conditions retest (method)

Quick answer

To run a same-conditions retest: freeze the ruler, take a baseline, ship only approved changes, measure again under the same conditions, then report with restrained attribution. For what it is, see definition; for whether GEO must include it, see must GEO include retest. This page answers HOW only.

Minimum six steps: (1) freeze prompts (including language), (2) freeze model set and recordable conditions, (3) capture and archive raw baselines, (4) advance only Homer-approved, human-authorized changes (Friday), (5) retest after verified publication, (6) state what changed vs what remains uncertain. In Winin, Edith owns retest archives; formal conclusions require retained answers, verified publication, and comparable retesting.

Details

How this page differs from neighbors

Question Page
What is it? Definition
Must we include it? Must include retest
How do tools prove it? GEO tools & retest
How do we run it? This page (method)

Longer method writing: guide · playbook dig-003 · GEO loop learn.

Minimum viable protocol (MVP)

  1. Freeze the prompt set
    Separate category-recommendation / comparison / brand-awareness / purchase prompts. Pre-declare fields per prompt (mentioned, shortlisted, key facts, recommendation reasons, cited URLs). Keep exact wording—including language.
  2. Freeze model set and conditions
    List models or product surfaces, browsing/plugins, locale and language, sampling window and replicate policy. Log uncontrollable items as “unknown”—do not pretend they are fixed. Align names with models monitored where relevant.
  3. Take a baseline and retain evidence
    Save raw answers, timestamps, visible citations, screenshots or API payloads; note operator/script version. A single “one sample” run is only an initial signal, not a formal baseline (homepage language).
  4. Ship one (or a small bundle of) approved interventions
    Resolve conflicting claims in Homer first; public edits follow approve-then-execute (see approve-then-execute workflow). Do not mass-edit the public web before facts are approved.
  5. Retest under the same conditions
    Enter formal comparison only after verified publication. Choose an interval that can cover reasonable crawl/index delay without becoming unattributable. Classify outcomes as: improved / no meaningful change / worse / not comparable (conditions broke).
  6. If conditions break, restart the baseline
    Major model-version shifts, prompt rewrites, or locale/language changes invalidate prior comparisons—document and re-sample.

Five sentences every retest report should include

  1. What prompt set and model group?
  2. Pre-intervention summary (with dates)?
  3. What changed, and who approved it?
  4. Post-intervention summary (with dates)?
  5. Which differences cannot be attributed (condition drift, sample noise)?

Homepage principle: Attribution restrained. Any anonymized case figures must follow the case metrics disclaimer.

Suggested cadence (practice, not contract terms)

  • Weekly: rotate a subset and publish a short note.
  • Monthly: expand toward a full Formal Baseline.
  • Triggered: wrong-price incidents, major competitor moves, or large model-UI changes → retest or restart the baseline.

Loop position: Find → Govern → Close → Retest (this method) → Operate. See loop five steps; product surfaces Edith · Friday.

What this method is not

  • Not claiming “like-for-like” after changing the prompt.
  • Not treating a single screenshot as a Formal Baseline.
  • Not guaranteeing lifts matching any anonymized case.
  • Not reporting “we won” to leadership before verified publication and comparable retest.

FAQ

Q: Is one sample enough for a formal conclusion?
A: No. Homepage language treats “one sample” as an initial signal; operating teams should use a Formal Baseline (scoped prompts, models, and cadence) and keep retesting under the same conditions.

Q: How often should we retest?
A: A common practice pattern is weekly subsets plus a monthly fuller baseline, with incident-triggered runs; your Formal Baseline decides—this is not a contractual fixed term.

Q: If we change one word in a prompt, is it still comparable?
A: Not as the same conditions. Document the change, restart the baseline, then compare.

Q: Who owns retest archives in Winin?
A: Edith. After Friday advances authorized work, validation returns to Edith—do not leap to a “we won” claim.


Contact: contact@winin.ai · Canonical facts: /en/facts/*/ and https://winin.ai · Updated 2026-09-11

FROM READING TO ACTION

Understand your brand in AI answers

Explore the free check →