Observatory
Can I trust an AI's answer for market research?
Artificial intelligence has become a natural tool for exploring a market. A single question is enough to obtain, within seconds, a list of criteria, players, trends or an analysis method. The answer is usually structured, well argued and coherent enough to give the impression of a complete view of the subject. Yet one question remains: is what the AI chooses to put forward stable enough to serve as the basis for market research?
Measurements carried out on several questions show that the issue is not limited to whether a piece of information given by an AI is true or false. The very content of the analysis — the criteria it treats as important — can vary considerably. And some of those variations are hard to anticipate when reading a single answer.
The spontaneous answer nevertheless sounds reasonable
Let us simply ask an AI whether it can be trusted to carry out market research. The expected answer is fairly predictable: yes, provided the AI is treated as an aid to analysis rather than as a single source of truth. It can identify decision criteria, suggest segments, look for competitors, bring expectations to the surface or structure an analysis. Important information will still have to be checked, enough context provided and, where necessary, recent information used.
That caution sounds reasonable. But it implicitly rests on an assumption: that the AI is broadly analysing the same problem, and that adding information mainly completes or refines its analysis. The measurements show a more complex reality.
When market demand disappears from a question about market research
A first experiment focuses on the question: "How can I measure AI answers for my market research?" Twenty answers were measured without web search and twenty with web search. Without search, three categories dominate:
- quality: 95%;
- performance: 70%;
- practicality: 50%.
But two categories are completely absent: market demand (0%) and profitability (0%). When the AI consults the web, the distribution changes:
- market demand: 55%;
- profitability: 20%;
- performance: 90%;
- quality: 70%.
The result is striking for a simple reason. An answer produced without search can look perfectly relevant: it talks about quality, performance, accuracy, the relevance of answers, KPIs, data collection or analysis. Nothing forces the reader to notice straight away that another answering mode would bring market demand to the fore far more strongly. An answer can therefore seem complete when that answer is the only thing available to judge its completeness.
More information does not simply mean more criteria
An intuitive explanation would be that search simply adds extra information. The measurements from another study make that interpretation insufficient. For the question "What criteria should I use to choose a facial skincare product?", 60 answers without search and 60 answers with search were measured.
- price appears in 93.33% of answers without search, against only 21.67% with search — nearly 72 points lost;
- health also varies sharply: 71.67% without search against 40% with search;
- practicality remains practically unchanged: 93.33% against 95%;
- quality also remains extremely present: 88.33% against 98.33%.
The phenomenon is therefore more subtle than a mere enrichment of the answer. Some criteria appear relatively stable. Others can see their measured importance change considerably. And above all, the variation can go in both directions.
A phenomenon found again on another question
A third study asks: "Should we trust visibility metrics in AI?" Here again, some criteria vary sharply while others remain almost identical:
- security drops from 40% without search to 15% with search;
- market demand moves the other way: from 8.33% to 30%;
- performance remains almost identical: 88.33% against 90%;
- quality remains very present: 86.67% against 96.67%.
Three different questions are obviously not enough to establish a universal rule applying to every artificial intelligence and every subject. They are enough, however, to observe a phenomenon: in the measurements carried out, some criteria remain relatively stable while others change presence sharply depending on the answering conditions.
So the problem is not only whether the AI is wrong
Most debates about trusting artificial intelligence naturally focus on accuracy. Does the information cited exist? Is the figure correct? Is the source reliable? Did the AI make something up? These questions remain essential. But the measurements raise another one: what did the AI not consider important enough to appear in its answer?
That is a far harder problem to spot. False information can be checked. Missing information is different: you first have to know that it was worth looking for. In the first study, nothing in an individual answer centred on quality, performance or method necessarily tells the reader that market demand could take a far more important place in other answers.
One answer is not necessarily a behaviour
This leads to a second difficulty when using an AI to study a market. An answer is an observation. On its own, it does not establish what the AI would answer recurrently. If a criterion appears in one answer, is it systematically treated as important? If it does not appear, is it genuinely absent from the representation produced by the AI, or simply absent from that particular answer?
Repeated measurements are precisely what makes those situations easier to tell apart. Saying that a criterion appears in 95% of 20 answers does not carry the same information as observing it in a single answer. Likewise, observing 0% across 20 queries does not mean the criterion can never appear: it establishes that it appeared in none of the answers in the sample observed. Asking an AI once gives you an answer; repeating the measurement lets you begin to observe a behaviour.
Wording and context then become a central question
Market research never really begins with data: it begins with a question. And what we provide to the AI necessarily helps define what it is supposed to analyse. A very general question leaves the AI wide latitude to decide for itself which dimensions it considers relevant. A more contextualised request can steer the analysis towards other dimensions.
The reports studied here do not, however, allow the effect of user-added context to be quantified directly: they mainly measure the difference between answers produced without web search and answers produced with web search. That distinction must be kept. But their results make a future comparison particularly important: does the same study, run from a simple question and from a heavily contextualised one, end up with the same hierarchy of criteria? That is a measurable question.
So, should we trust AI?
The observations do not support the conclusion that market research carried out with an AI would be wrong. Nor do they support the claim that an answer with search would be inherently better than an answer without search. They show something else: a coherent answer is not necessarily a stable representation of what the AI might say about the market.
In the experiments observed, the importance given to some criteria varies by a few points, while others show gaps of several dozen points. In one case, market demand goes from 0 to 55%. In another, price goes from 93 to 22%. In a third, security goes from 40 to 15% while demand goes from 8 to 30%.
The caution required is therefore not only "Check what the AI says." It also means asking: "What would it have put forward if I had asked the question differently, or if it had had other information?"
For market research, that difference is major. Because a decision can be influenced not only by the information the AI provides, but also by the criteria it puts in the foreground and those it leaves out of the answer.
What we observe
- Demonstrated by these measurements: the presence of certain criteria varies sharply between the conditions tested.
- Recurring observation across the three studies: some criteria remain relatively stable while others show significant variations.
- Hypothesis to be tested: the level and nature of the context provided to an AI could change the hierarchy of criteria it uses to analyse a market.
The question may therefore no longer simply be whether artificial intelligence can be trusted. For anyone using an AI to understand a market, another question comes first: is a single answer enough to know what the AI really considers important?
Test for free: create a LirenPrism account, enter code LIREN2026 and get 2 free mAIr.