First define the surface
| Surface | What it observes | Key limitation |
|---|---|---|
| Closed-book API | Model knowledge without a live search layer | Does not represent the provider's consumer search product |
| API with retrieval | A configured model and retrieval toolchain | Results depend on the selected search, tools and implementation |
| Consumer AI product | The experience customers may actually use | Account, location, personalization and product changes can affect output |
| Search AI feature | Generated answers grounded in a search index | Eligibility, indexing and result presentation vary by market and query |
Do not combine these surfaces into one trend line. Record the provider, product, model or version when visible, search state, account mode, language, location and timestamp.
Build questions from customer decisions
Identity
Who is the brand and what does it do?
Category discovery
Which providers or products fit a category without naming the target?
Scenario fit
What is suitable for a specific buyer, task or constraint?
Comparison
How does the target compare with named alternatives?
Evidence
Which sources support the important claims?
Risk boundary
When is the product or service not appropriate?
Brand-named questions measure understanding. Non-branded category and scenario questions measure discovery or shortlist opportunity. They need separate denominators.
International model coverage
English-language programs prioritize OpenAI, Claude and Gemini surfaces when they match the target market and buyer behavior. Coverage follows the customer's real market, not a fixed logo checklist.
Each run records the provider, model or product surface, retrieval condition, language, market and time. A translated question is not automatically an equivalent observation because terminology, available sources, search indexes and product features can differ.
The minimum observation record
- Frozen question family, exact wording and version hash
- Provider, product surface, model or version when disclosed
- Retrieval or search state, account mode, language, location and time
- Complete raw answer and visible source URLs
- Failure, refusal, rate-limit and missing-citation states
- Brand mention, factual support, shortlist, recommendation and citation judged separately
- Run identifier and repeated sample index
Read the result as a distribution
Compare models by question family and repeated observations, not by one screenshot. Show the numerator and denominator behind every rate. When the question set, model, search mode or scoring rule changes, create a breakpoint instead of drawing one continuous trend.
A useful monitoring cycle ends with an action decision: continue observing, correct a fact, improve a page, strengthen a source or investigate a new issue. Monitoring without an operating response becomes a dashboard rather than a management system.
Continue with the measurement method
See how to measure GEO performance and the Winin measurement method for denominator, baseline and post-test requirements.