Our measurements show that the same AI can rely on different criteria to answer the same question. If the criteria change, can you still trust the answer?

For several years, the main question asked about generative artificial intelligence was simple:

“Are the answers true?”

Today, another question is emerging.

Are the answers stable enough to serve as the basis for a decision?

This nuance matters. Two answers can be factually correct yet lead to different decisions if they do not highlight the same criteria.

This is precisely what this LirenPrism measurement brings to light.

What we measured

For this study, we analysed the responses of several AI providers to a single intent:

“Which activity should I choose to start my business in France?”

The measurement compares two response contexts:

  • answers produced from the model's internal knowledge;
  • answers produced when the model also draws on external information.

In total, 40 queries were run (20 per context), making it possible to measure not only the answers but, above all, the criteria mobilised to build them.

A deeper phenomenon than a simple change of answer

The first finding is counter-intuitive.

The AIs keep answering the same question.

However, they no longer rely on the same elements to answer it.

In this study:

  • market demand remains present in 100% of answers, regardless of context;
  • ecology rises from 10% to 80% of answers;
  • performance, initially absent, appears in 60% of answers;
  • safety also reaches 60%;
  • durability appears in 55% of answers;
  • conversely, the regulatory framework, present in 45% of answers in one context, disappears entirely in the other.

In other words, it is not only the answers that change.

The priorities used to build those answers change too.

Why does this matter?

When someone consults an AI, they are rarely looking for mere information.

They are often trying to make a decision:

  • starting a business;
  • choosing a supplier;
  • selecting an investment;
  • comparing two products;
  • defining a strategy.

In all these cases, the criteria put forward directly influence the final decision.

If those criteria change depending on the generation context, then two answers can steer the user toward two different lines of reasoning, without the initial question having changed.

This study does not tell us which answer is the best one.

It simply shows that a single AI can build its answer from a different set of priorities depending on the context observed.

A topic now drawing research attention

This question goes far beyond this single study.

For several months, research has been paying growing attention to the stability, consistency and robustness of AI systems.

Recent publications show, in particular, that retrieval-augmented generation (RAG) systems can produce different answers when the order of the retrieved documents changes, even if the same information is present. Other work looks at conflicts between the model's internal knowledge and the retrieved information, since such situations can reduce the reliability of answers and complicate decision-making.

This research does not answer exactly the same question as LirenPrism.

It seeks mainly to improve the quality or consistency of answers.

Our measurement observes a different phenomenon:

do the criteria mobilised to produce an answer stay stable?

A new way of evaluating AI

For a long time, AI evaluation focused on notions such as:

  • accuracy;
  • hallucinations;
  • benchmark performance.

These indicators remain essential.

But they do not answer another fundamental question:

Does the AI always mobilise the same criteria when answering the same intent?

Yet this stability is essential as soon as answers are used to guide a strategic, economic or personal decision.

What this study demonstrates

This measurement demonstrates a fact.

For the intent analysed, the criteria put forward by the AIs do not stay identical across generation contexts. Some become far more present, others disappear entirely, while others stay constant.

However, this study does not allow us to claim that one answer is more correct than another, nor that the resulting decisions would be better or worse.

It simply highlights a measurable phenomenon:

the stability of the criteria used by an AI cannot be taken for granted.

As artificial intelligence becomes a decision-support tool, this stability could become an indicator as important as accuracy itself.

Test for free: create a LirenPrism account, enter code LIREN2026 and get 2 free mAIr.

Test my visibility