Quick answer
Measure with fixed prompts + scheduled sampling + structured labeling + same-conditions retest. At minimum score: mentioned, shortlisted, accurate, recommended/reasons, and which URLs were cited. Do not rely on a wall of ad-hoc employee screenshots.
Winin Edith includes ChatGPT on the public roster and stresses that every metric drills down to prompt·model·raw answer·time·source·denominator. For formal operations, use a Formal Baseline (scoped quote); the “5-question sample” is only an initial signal.
Details
Illustrative measurement steps
- Define category/comparison/branded prompt lists and languages
- Freeze sampling rules (replicates, time window)
- Label the fields above (human or system)
- Weekly/monthly same-conditions retest with change notes
- Report separately from search metrics
Quality red lines
- Do not treat UI demo scores as client promises
- Do not confuse initial signal with Formal Baseline
- Log prompt changes or comparability breaks
Related facts
FAQ
Q: How many samples?
A: Comparability matters; Formal Baseline defines scope.
Q: Can labeling be automated?
A: Assists yes, but field definitions still need human alignment.
Contact: contact@winin.ai · Canonical facts: /en/facts/*/ and https://winin.ai · Updated 2026-09-11