← Monitoring & Diagnosis

Guides

How to monitor AI recommendations: from question set to comparable retest

Same-conditions retest means measuring again with the same questions under the same conditions, so that before and after can be compared, while stating clearly which part changed and which part remains uncertain. Its value is not that you measured many times, but that the two numbers can be read side by side.

1. Why a baseline comes first

AI answers change. The same question may get a different answer today than next week: model versions update, retrieval results shift, competitors publish. Without a fixed baseline and complete archives, you cannot tell management whether share or reasons improved after publishing fact pages, and you cannot separate "the change came from our action" from "the change came from model drift". Without a baseline, the discussion turns into a contest of opinions.

2. Step one: turn the question set into a reusable asset

Start from what buyers would actually ask, not from keywords.

Example: suppose you sell an enterprise project management SaaS. Prompts could be "what project management tool should a small team with a tight budget use", or "is there project management software that supports on-premise deployment". These are different intents with typically different candidate sets, so record them separately.

Rules:

  1. Each prompt has a fixed language, fixed wording and fixed punctuation. Chinese and English are two different prompts.
  2. Start with roughly 10 to 30 prompts covering category recommendation, head-to-head comparison, usage scenarios, alternatives and budget constraints. Do not compress several intents into one sentence.
  3. Freeze prompts once recording begins. Add a new prompt rather than editing an old one in place.
  4. Number each prompt and add one line describing the buying intent it represents, so others can follow your reasoning.

3. Step two: fix everything that can vary

  • Model group: for example one global group and one Chinese-language group. Write the list down and do not swap members mid-stream.
  • Entry point and switches: consumer product or direct API, retrieval on or off, logged-in or personalised memory state. Fix all of them.
  • Sampling window: record date and time window, and state how many samples per prompt.
  • Archiving: keep the full raw answer, not just a yes/no tick.

4. Step three: record in one shared table

Suggested columns:

date | model | query | language | entry_retrieval | mentioned | shortlisted | recommended | reason_owner | accuracy_notes | cited_urls | competitor_set | uncertainty_notes

reason_owner records who the recommendation reasons belong to, and uncertainty_notes records the part of this result you cannot explain. Both are heavily used later when you do attribution.

5. Step four: compute five numbers separately, with the denominator stated

Mentions, recommendations, citations, visits and opportunities must be counted separately, and each number must trace back to a denominator.

Example: suppose you use 10 models x 20 prompts x 1 sample each, giving 200 "answer slots". The result: 60 mentions, 18 appearances in the recommendation list, 5 citations of your pages. These three cannot be summarised as a single percentage, and "60 mentions" cannot be used to explain business opportunities. Visits and opportunities need separate measurement (parameterised links or channel-level reporting, for instance), and you still have to state that other explanations exist. All of the above numbers are illustrative, not measured results.

6. Step five: retest only after a published change

A retest presupposes that you actually published a checkable, citable factual update. Record the timestamp and URL of the change so the retest can be matched to it. If you only revised an internal document, the outside world cannot see it and the retest means nothing.

7. Write the three layers into the proposal

  • Initial signal: a small set of fixed questions, one sample each, a few model interfaces. Its role is to show direction, and it is explicitly not a formal baseline.
  • Formal baseline: custom prompts, models and cadence; a full baseline that can be retested repeatedly.
  • Brand and product GEO execution: strategy plus site, content and channel execution, validated by returning to same-conditions retest.

Spelling out these layers mainly prevents using initial-signal results to promise execution-layer outcomes.

8. Suggested cadence (practical reference, not contractual)

  • Weekly: rotate a subset for retest, keeping continuity.
  • Monthly: run the full baseline.
  • Triggered: retest immediately after a major competitor launch, or after factual errors such as wrong pricing or discontinued status.

9. How to word the results

Correct: "Under identical prompts and the same model group, the proportion entering the recommendation list moved from A to B; during that period we published a factual page; some models still do not cite it, for reasons not fully determined."

Incorrect: "Because we did GEO, the improvement is definitely ours."

The first wording keeps uncertainty visible; the second treats correlation as causation.

10. When conclusions are unusable

  • Drawing a verdict from a single sample. One sample is a signal only.
  • Calling something year-over-year after changing the prompts.
  • Treating a model's own version change as your optimisation result.
  • Using the uplift in someone else's anonymous case as your own expectation.
  • Treating one platform's statement about crawling and indexing as evidence that every model will recommend you; different platforms work differently.

FAQ

Is one sample per question enough?

As an initial signal, yes, provided you state clearly that it is not a formal baseline. Operational judgement needs a fixed baseline and repeated sampling.

How should the denominator be written?

Write it in a form anyone can recompute, for example "number of models x number of prompts x number of samples". Every metric must trace to that denominator, otherwise movement cannot be explained.

Winin's approach here is to keep monitoring, fact governance, authorised execution and same-conditions retest in one workflow: get a comparable baseline first, then decide what to change, then confirm the change under identical conditions afterwards. Scope depends on your question set and model list.

FROM READING TO ACTION

Understand your brand in AI answers

Explore the free check →