Measure four separate things.Whether the brand is mentioned, whether the description matches confirmed facts, whether the brand enters the recommended set, and whether the answer carries a verifiable source. Each has its own denominator. Collapsing them into one number makes a score change impossible to act on.
- Mentioned is not recommended, and recommended is not accurate.
- Every metric needs a stated denominator: how many questions, how many samples.
- Questions must be frozen and versioned, or two runs are not comparable.
- When the denominator is thin, publish no score rather than a zero.
- Closed-book and retrieval-augmented answers are separate tracks.
Why a single score fails
A score moving from 62 to 55 tells you nothing about what to fix. Worse, mentions can rise while accuracy falls — the model talks about you more often and gets you wrong more often — and a single number reports that as an improvement.
The denominator matters more than the numerator
Appearing twelve times means nothing without the number of questions, repeats per question, engines covered and the time window. When that base is too thin, the honest output is no score.
Frequently asked questions
How long before GEO results are visible?
There is no reliable public figure for this: it depends on re-crawl and index cadences that engines do not publish. Rather than accepting a promised timeline, require a pre/post protocol and a denominator so the timeline is measured.
Can keyword rankings substitute for GEO metrics?
No. Generative engines do not return ranked lists, and the same question yields different answers across runs, so the unit of measurement is a sampled rate, not a position.