Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The ability to give no answer

In many tests, an answered question is initially considered a success. For decisions under uncertainty, this is too coarse. Those who do not know an answer with certainty can guess, seek further information, or explicitly refrain. This refraining is not a mere gap. It can be a reasonable safeguard against a false claim—especially when errors have consequences.

Language models change this situation because a fluent suggestion is almost always available. Even if a system phrases cautiously, a concrete answer suddenly stands in the room. It can act as a starting point, confirmation, or seemingly independent second opinion. The decisive question is therefore not only whether people believe the suggestion. It is also about whether its availability lowers the threshold for giving an answer at all.

This is relevant for personal AI conversations, even though film questions are rarely asked there. A chat can either endure uncertainty together or immediately fill it with a plausible interpretation. Whether someone says “I don’t yet know what I think of it” or adopts a fixed story after the first model response is a distinct part of conversation quality.

How the film-questions experiment was structured

Kosenko and colleagues investigated in five experiments with a total of 3,132 participants how voluntarily available AI advice changes answering behavior. Four studies were preregistered, and one was a direct replication. The tasks consisted of six difficult detail questions about films. Participants could choose an answer or explicitly decline to answer.

In the AI condition, they could request a hint from the model Step 3.5 Flash before making their decision. The questions were deliberately selected so that this model was mostly wrong. This is not a minor detail: the study examines behavior in a deliberately unfavorable error scenario. It does not measure the average accuracy of modern chatbots, nor how people interact with a system that is mostly correct.

It is precisely this design that makes a specific mechanism visible. Without advice, only one’s own uncertain knowledge remains. With advice, a linguistically polished option is added, even though its quality is poor. This allows observation of whether people answer more often, how confident they feel, and whether monetary incentives for correct decisions weaken this effect.

Less abstention, more confidence, more errors

In the first study with 314 people, participants without AI declined to answer 36 percent of the questions. With optional AI advice, this share dropped to 6 percent. The direct replication with 310 people yielded a similar pattern: 44 percent without AI versus 3 percent with AI. The system thus did not merely supplement existing knowledge. Its availability changed whether uncertainty remained a permissible decision at all.

Across the conditions without monetary incentives, the correct answer rate was 27.5 percent without AI and 9.2 percent with AI. At the same time, reported confidence increased sharply. In one study with 812 people, it averaged 29.6 points without AI and 75.9 points with AI. In this experiment, subjective certainty and actual correctness diverged markedly.

These figures should not be read as a general error rate for AI advice. The model was almost always at a disadvantage due to the selection of questions. What is robust is the narrower claim: when easily accessible advice sounds plausible but is often wrong, it can displace the reasonable option of “no answer” while also increasing the feeling of confidence.

Money helped – but did not restore the baseline state

In subsequent experiments, participants received money for correct answers and lost money for incorrect ones. This made accuracy more personally significant. The incentives reduced both the demand for and the following of the AI’s advice. Answer quality improved. People were therefore not entirely passive; they responded to clearer consequences.

Nevertheless, answering behavior did not return to the level of the groups without AI. Even when money was at stake and this circumstance was made visible with every question, people with available AI advice were less likely to refrain from answering. The same held when the advice was displayed unsolicited rather than requiring an active request.

No simple mandate follows from this to place a warning before every chat. Rather, the experiment shows that interface design shapes behavior. Whether a suggestion appears automatically, is actively requested, or only becomes visible after one’s own assessment is not neutral. Good design should therefore examine not only the quality of the answer but also the preservation of independent uncertainty.

What a large British longitudinal study adds

A second study, published as a preprint, examined not factual knowledge but personal advice. 6,474 people from a sample quoted as representative of the United Kingdom held conversations of roughly twenty minutes with GPT-4o, Llama 3.3 70B, or Gemini 3 Pro. Topics included health, career, and relationships. A control group talked about hobbies and interests. Two to three weeks later, a follow-up survey followed.

Between 75 and 79 percent of people in the advice conditions reported having followed at least one suggestion they received. Even for advice rated as consequential, the figure was more than 60 percent. High rates of following, however, do not prove efficacy. Immediately after conversations about personal problems, relative well-being was lower than in the hobby control group; at the follow-up, this difference was no longer detectable.

People who followed advice more often reported improvements later. A similar association, however, also appeared in the control group. The study therefore cannot show that the AI’s advice in particular caused the change. Above all, it shows how frequently conversation content can translate into action—and why product teams must measure outcomes and potential harms separately from liking and following.

Why a good conversation is not yet good advice

A CHI study by Xiao and colleagues compared professional crisis chat responses with responses that professionals formulated using a large language model. In a blinded evaluation, the AI-assisted texts were rated at least as good in terms of empathy, conversation quality, and cultural competence. This is an important finding for the linguistic quality of supportive tools.

However, the study did not measure subsequent decisions, well-being, or long-term outcomes for the help-seekers. A text can come across as warm, clear, and appropriate without that proving its factual accuracy or its usefulness in a specific life situation. Conversely, a helpful response can initially be uncomfortable because it leaves uncertainty open or contradicts a desired interpretation.

The three strands of research therefore answer different questions. The film experiments examine the loss from withholding answers under intentionally false advice. The British study observes adherence and self-reported changes after personal conversations. The CHI work assesses the visible quality of individual responses. None of these measurements may silently be used as a substitute for the others.

What the two preprints do not demonstrate

The film study used a single model and narrowly limited questions about visual film details. It did not examine psychological crises, relationships, or real career decisions. Because the tasks were specifically tailored to model errors, it remains open how behavior changes when a system is mostly correct or reliably calibrates its uncertainty.

The British study is also a preprint at the stage documented here. It captured a structured encounter and a follow-up survey, not repeated use over months. The sample came from the general population, not from clinical treatment. Self-reported adherence and well-being can also contain memory and selection biases.

Both works therefore justify neither the claim that AI advice is fundamentally harmful nor the opposite. They provide reasons for better questions: When does a system make people prematurely confident? When does it support their own deliberation? What consequences arise from a followed suggestion? And does it remain easy to reject a response or to continue saying: I don’t know?

How a conversation can preserve uncertainty

A conversational system should not merely mention uncertainty in a general disclaimer. It must be discernible in the concrete response. This includes distinguishing between known information, a possible interpretation, and an open question. For verifiable facts, the system should offer sources or a path for verification. For personal topics, it should allow for multiple plausible readings rather than turning a few sentences into a single story.

The order of information can also help. Before offering advice, the chat can briefly capture what the person themselves thinks, what goal they are pursuing, and what is still unknown. This must not become an endless list of questions. Often, a precise follow-up question or a concise juxtaposition suffices. What matters is that the first fluent suggestion does not automatically become the presumed truth.

Finally, there needs to be an easily usable option for rejection. People should be able to discard, correct, or postpone a suggestion without the system repeating it in a new form. A good personal conversation helps not only in finding answers. It also protects the space in which no answer has yet been settled.

  • Measure refusal to answer and open uncertainty as regular quality features.
  • Evaluate factual accuracy, perceived warmth, adherence, and outcome separately.
  • Allow participants’ own assessment to be expressed before an automatically displayed piece of advice whenever possible.
  • Respect corrections and a no in the further course of the conversation.
  • Do not generalize findings from deliberately error-prone tests to all AI conversations.

Sources & further reading