Observatory
What if the AI answering you wasn't always the same?
Two modes, two lines of reasoning, two possible answers. Depending on whether an AI draws on its internal memory or consults the Web, the criteria it foregrounds can change. Our measurements show this on a specific question — even though the user cannot see the difference.
For a long time, we treated an AI answer as a single block. You ask a question, the AI answers, end of story.
Yet one question is rarely asked: how was that answer constructed? Was it produced from the model's internal knowledge, or after consulting the Web?
The difference seems subtle. Our measurements show it can profoundly reshape the resulting discourse.
One question. Two different lines of reasoning.
For one and the same question — “Should we trust visibility measurements in AI?” — we compared two situations:
- the AI answers using only its internal knowledge;
- the AI answers after consulting the Web.
The result is not simply a longer or more recent answer. The criteria it foregrounds change too.
For instance, on this measurement:
- the security criterion is markedly more present when the AI answers without consulting the Web;
- conversely, market demand becomes far more prominent when web search is enabled.
In other words: the AI does not merely add information. It can reorganize the criteria it treats as important.
One answer — but which engine actually did the work?
To the user, the two answers look alike. In both cases the style is similar, the formatting is similar, and the apparent level of confidence is identical.
Yet the reasoning that led to the answer is not necessarily the same. That is an important difference.
Today, large models alternate between internal knowledge and external information retrieval depending on the product's capabilities, the available settings, or the nature of the question. This architecture is widely known as Retrieval-Augmented Generation (RAG), which combines the model's knowledge with information retrieved from external sources.
But in most interfaces, the user does not always know precisely which mode was used to produce the answer.
This is not only a technical question
This distinction becomes strategic. Two decision-makers can ask exactly the same question. One gets an answer based mainly on the model's memory; the other, an answer enriched by the content currently available on the Web.
Both answers can be relevant. Yet they can steer the decision in different directions.
Our measurements do not explain why this difference appears. They simply show that it exists for the question studied.
Why this changes how we interpret an AI
For years, we have evaluated AI by comparing their answers. Tomorrow, we will probably need to start by understanding how those answers were produced.
Because two seemingly identical answers can rest on two different constructions:
- one drawn mainly from the model's memory;
- the other influenced by the information the Web makes visible at the moment of the query.
Recent work also shows that generative search systems differ significantly in their reliance on internal knowledge, on external sources, and in the stability of their answers — introducing new dimensions of evaluation beyond the mere quality of the final answer.
A new question for businesses
When a company seeks to improve its visibility in AI, one question becomes essential: is it trying to influence what the AI “knows”… or what the AI “finds”?
These are probably not the same levers. And they may not call for the same strategies of communication, content, or web presence.
What this study shows
Established fact
For the question analyzed, enabling web search comes with measurable changes in the criteria foregrounded by the AI — notably security, market demand, and quality.
Recurring observation
Modern generative search architectures increasingly combine models' internal knowledge with information retrieved in real time, which can alter both the content and the structure of answers.
Open question
In the future, will users systematically be able to know whether an answer comes mainly from the model's memory or from a web search? This study cannot answer that. It only shows that this distinction can have a measurable impact on the discourse produced.
Test for free: create a LirenPrism account, enter code LIREN2026 and get 2 free mAIr.