Quick answer
Without web search (closed book), a model answers only from what it learned in training; with web search, it searches the web first and answers from what it finds. The same question can produce completely different brand recommendations in the two modes. Both are useful for measuring brand performance, but they must be counted separately: web-search results reflect roughly what users see in an app with search on, while closed-book results reflect the model's "memory" of your brand.
Details
Two ways of answering
| Closed book (no web search) | Web search | |
|---|---|---|
| Source of information | Training data, with a cutoff date | Web pages retrieved at answer time |
| Shows sources? | Usually not | Usually attaches citations or a source list |
| Influenced by | Model version, training corpus | Search results, page content, ranking |
| How fast it changes | With model version updates | As soon as web pages change |
| Best for understanding | The model's baseline impression of a brand | What users actually see recommended today |
Why answers differ between the API, the web version and the app
The same model can answer differently depending on the entry point, usually for four reasons:
- Web search on or off: app users may have web search on, while APIs often default to closed book and need a separate retrieval parameter.
- System prompts and parameters: system settings, temperature and output length differ by entry point.
- Model version: the web version may have been updated while the API still uses a pinned version.
- Personalisation and context: the app may carry the user's chat history or location.
So before saying "DeepSeek recommended X", state which entry point, whether web search was on, the exact question, and how many times it was asked.
Search queries and query fan-out
When answering with web search, a model usually does not search with the user's exact words. It splits the question into several search queries; for example, "home coffee machine recommendations under a 3,000-yuan budget" might become "home coffee machine recommendations" and "coffee machine reviews under 3,000 yuan". This is often called query fan-out.
For brands, this means whether you are found depends on these derived search queries, not only on the user's wording. Some platforms' APIs return the queries actually used; record them when available. When they are not visible, do not guess, and do not treat your guesses as the platform's real behaviour.
Should you turn on web search when measuring brand performance?
Measure both, and count them separately:
- Web-search results: the cited sources show directly whose content AI reads, which guides content and channel work.
- Closed-book results: show the model's baseline knowledge of the brand and reveal outdated or wrong "memories".
Never combine the two into one recommendation rate. Otherwise, when the metric changes you cannot tell whether web content made the difference or the model version changed.
Common questions
Q: Are API test answers the same as in the app?
A: Not necessarily. Entry point, web-search state, parameters and version may all differ. Reports should state the entry point and conditions before any comparison.
Q: What is query fan-out?
A: When answering with web search, the model splits one question into several search queries. Whether a brand is found depends on those queries.
Q: Should web search be on when testing AI recommendations?
A: Yes, but test closed book as well, count the two separately, and never merge them into one recommendation rate.
Related facts
Contact: contact@winin.ai