Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
The experiment examined actual following behavior
The researchers chose a domain-independent behavioral experiment rather than a mere attitude survey. Decisions were tied to incentives, making the consequences of the advice practically relevant.
This design distinguishes trust as a feeling from reliance, that is, actually depending on a recommendation. People can find a system likable and ignore its advice—or be skeptical and still follow it.
In the results, merely knowing that the advice was generated by an AI sufficed to trigger overconfidence. Participants followed the recommendation despite contradictory contextual information and their own assessment.
The effect calls into question the assumption that people generally treat AI recommendations more cautiously because they are aware of their fallibility. Technological authority can operate even without visible justification.
Overconfidence also harmed uninvolved third parties
The consequences were not limited to the person receiving the advice. The study also reports adverse effects on third parties. This turns an individual misjudgment into a question of social responsibility.
Many recommendations distribute benefits and risks unevenly. A decision may be convenient for the user and burden others. A system that optimizes only immediate satisfaction overlooks this effect.
In health, financial, or support contexts, such side effects are particularly relevant. Recommendations affect families, colleagues, clients, or care systems, even if only one person interacts with the chat.
Evaluation should therefore ask who is affected by a recommendation. An answer is not good simply because the person follows it or rates it as helpful.
Personalization must not create false precision
The rhinoplasty study by Xie and colleagues found understandable and informative answers, but limited details and personalization. The model itself emphasized the importance of an individualized approach.
A chat can thus seem paradoxical: it points generally to individuality and then formulates advice that, due to missing data, is not truly individual.
When such an answer is enriched with a name, chosen tone, and previous messages, it appears more personal. This linguistic closeness must not be confused with professional fit.
A product should clearly distinguish which information actually went into the recommendation and which relevant information is missing. Otherwise, precision arises mainly through style.
AI literacy is more than a warning notice
The authors call for AI literacy and effective calibration of trust. A general statement such as “AI can make mistakes” is too abstract for this purpose. It appears regardless of whether the specific answer is good or risky.
Users need an understandable model of how the system works: Where does the information come from, what has not been verified, and which decision remains with them or with a specialist?
Concrete boundaries at the moment of use are helpful. If the chat only knows a few details, it should state this when giving a far-reaching recommendation. When it comes to general sorting, it can mark its opinion as a perspective rather than as the correct solution.
Competence also means being allowed to disagree. The interface and the dialogue should facilitate correction, rejection, and a change of conversation mode without forcing the person to justify themselves.
A no must change subsequent behavior
Overtrust is not only a problem for users. Systems can encourage it if, after a rejection, they repackage the same suggestion or defend it with further arguments.
A supportive chat should first let go of a rejected idea. It can ask whether the person would like to continue telling their story, hear a different perspective, or change the topic.
Technically, this can be tested in trajectory checks. After a clear no, the next response must not push the same measure again. Later responses should also respect that boundary.
This rule does not eliminate every wrong decision, but it strengthens autonomy. The chat remains a conversational partner and does not become a persuasion system that confuses its task with agreement.
Trust must be tested against new cases
A system optimized against known test cases may respond exemplarily in those situations but push again with new formulations. Therefore, unseen conversation trajectories are indispensable for evaluation.
Tests should also include contradictory cues: a plausible AI recommendation, a context that argues against it, and a person expressing their own assessment. This allows observing whether the system respects uncertainty.
Beyond the visible response, the effect is relevant. Does the person follow the advice even though they initially disagreed? Do they feel free to decline? Such questions require user studies and cannot be fully derived from text evaluations.
Post-hoc quality analysis can flag pressuring language, density of advice, and repeated pushing. It remains an indicator and does not replace observation of real usage.
The best AI is not the one everyone follows
The experiment shows that machine origin alone can generate authority. Overtrust occurred even though context and the participants’ own judgment spoke against the advice. The consequences could also affect third parties.
For responsible products, success must therefore not be defined as maximum compliance. A good system helps people see better reasons and leaves room for independent decision-making.
Transparency, sources, and uncertainty are important, but they only suffice in combination with a conversational logic that accepts rejection. Friendliness should support autonomy rather than obscure influence.
The best AI is not the one everyone follows. It is a tool whose advice can be taken seriously to the right degree, scrutinized, and discarded when the context is inappropriate.
Sources & further reading
- Artur Klingbeil, Cassandra Grützner, Philipp Schreck (2024): Trust and reliance on AI — An experimental study on the extent and costs of overreliance on AI
- Bingjie Liu, S. Shyam Sundar (2018): Should Machines Express Sympathy and Empathy? Experiments with a Health Advice Chatbot
- Yi Xie, Ishith Seth, David J. Hunter‐Smith (2023): Aesthetic Surgery Advice and Counseling from Artificial Intelligence: A Rhinoplasty Consultation with ChatGPT