Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The guiding question lies behind the conversation surface

In their review published in 2023, Coghlan and colleagues examine central ethical problems of AI-supported conversational systems in the field of mental health. Their starting point is deliberately ambivalent: such systems can promote access to information and offerings, but at the same time they bring risks that may be exacerbated in people with psychological distress. The publication addresses four important ethical problem areas using a recognized framework of five principles. The present source description does not name these four areas nor the individual principles; a more detailed enumeration would therefore be speculative.

What is decisive, in any case, is less the number of categories than the perspective of the work. Ethics does not appear as a subsequent control of individual responses, but as a task along the entire technical development and use. The authors accordingly address developers, providers, researchers, and professionals. This shifts the view from the seemingly private exchange between human and software to the organization behind the system: Who determines the intended use, what promises are made, how are risks monitored, and who can take action?

An ethics review is not evidence of efficacy

The lead publication is a critical review, not a clinical intervention study. It neither examines symptom trajectories in its own sample nor does it randomly compare an AI system with human treatment. Its results consist in the identification and ethical analysis of problems as well as in recommendations for responsible development and introduction. This is a relevant contribution, but it does not answer the question of whether a specific product is effective or safe for a specific group of people.

This methodological limitation is central to the public debate. From a convincingly described benefit, no proven effect can be derived; conversely, from an ethical problem description it does not follow that every application would be unjustifiable. In our editorial view, the two levels are too often conflated. Providers can point to good conversation experiences even though long-term consequences are unresolved. Critics can name real risks without knowing their frequency. A responsible assessment must therefore keep at least three questions separate: What does the system achieve technically, which consequences are empirically documented, and under what conditions is its use normatively justifiable?

More access can also mean more responsibility

The potential access advantage is not merely a marketing argument. Coghlan and colleagues explicitly acknowledge that conversational systems can improve access to information and services. Zhang and Wang additionally cite constant availability, scalability, lower geographic barriers, and the possible willingness to disclose more openly to a machine. The article presents such characteristics as potential responses to scarce professional resources and long waiting times.

But reach does not eliminate responsibility; it multiplies it. A system available at any time can also be misunderstood at any time, raise false expectations, or capture sensitive data. Its low access threshold does not automatically make quality differences recognizable to users. It becomes particularly problematic when ease of use is confused with professional reliability. Our assessment is therefore: access is an ethical gain only when the additional people reached are not simultaneously left alone with risks that have arisen through product design, marketing, or unclear responsibilities.

The comparison with human professionals narrows the debate

In 2024, Zhang and Wang raise the pointed question of whether AI could replace psychotherapists. The article, however, does not report its own randomized replacement trial. It consolidates research and application examples, describes potential advantages and limitations, and itself calls for large-scale randomized controlled trials. It mentions preliminary indications of a short-term reduction in anxiety and depression symptoms through AI-supported offerings. At the same time, the authors point out that the underlying studies often have small groups and no long-term follow-up, and that longer-term effects remain uncertain.

The replacement question is therefore empirically premature and conceptually too coarse. In this article, AI encompasses very different functions: prognosis, detection, support for professionals, monitoring, virtual environments, and text-based interventions. A system can take on a clearly limited task without replacing a professional role as a whole. Despite far-reaching expectations for the future, Zhang and Wang themselves arrive at a complementary position: human oversight, ethical judgment, and the ability to interpret complex emotional and nonverbal signals remain essential.

From an editorial perspective, we also consider the replacement comparison to be politically distracting. It directs attention to a hypothetical competition between human and machine, while more concrete decisions are already pending today: Which function may be automated? When must a human be reachable? Which claims about performance are permissible? Such role questions can be examined. The blanket question of complete replacement, by contrast, generates more speculation than guidance.

Emotional language is not yet a relationship

Zhang and Wang cite studies in which ChatGPT was able to recognize emotional components of hypothetical situations and reproduce them with linguistic nuance. The article specifically mentions an assessment using the Levels of Emotional Awareness Scale. Such results can show that a language model produces responses that appear emotionally attentive according to predefined criteria. The authors themselves emphasize, however, that this performance is based on pattern recognition and language modeling, not on experiential understanding.

This distinction is not a philosophical side issue. For users, the inner workings are barely visible in conversation: a fitting formulation can seem like sympathy, commitment, or a reminder of a shared history. But neither a reliable relationship nor clinical efficacy follows from this. Moreover, a benchmark with hypothetical scenarios measures something different from the course of real, longer-term conversations under stress. Perceived empathy can be a product feature; it is not evidence that the system can take responsibility or reliably grasp the meaning of a situation.

The ethical map is now broader

The scoping review by Meadi and colleagues from 2025 expands the earlier foundational analysis with a systematic mapping of the literature. Searches were conducted in seven scientific databases; additional texts were added via snowball search. English and Dutch publications of various types were considered, excluding symposium abstracts. Two researchers independently assessed eligibility. An ethical concern was identified as a separate topic if it appeared in more than two articles.

A total of 101 publications were included, 96 of them from 2018 or later. Reviews formed the largest group with 22 contributions, followed by 17 commentaries. The researchers distinguished ten topics. Privacy and confidentiality were addressed most frequently, namely in 62 of the 101 texts. Safety and harm appeared in 52, justice in 41, efficacy in 38, and responsibility or accountability in 31 contributions. Further topics were empathy and humanity, explainability and trust, anthropomorphization and deception, autonomy, and concerns about jobs in healthcare.

These numbers measure attention in the literature, not the frequency of real harms, nor their severity. A frequently discussed problem is not automatically the largest; a rarely addressed one can be considerably underestimated. Moreover, a considerable part of the material consisted of reviews and commentaries. The authors themselves identify an insufficient representation of stakeholder perspectives and under-researched areas. The overview is therefore a map of the debate, not a conclusive risk assessment.

Data protection is part of the depicted relationship

That privacy and confidentiality occur most frequently in the scoping review is consistent with the subject matter. Conversations about psychological distress can contain particularly sensitive information. The problem concerns not only abstract data security but also the significance of the conversational context: users may perceive a protected, personal exchange, while technically a data-processing service operates with its own storage, analysis, and access structures.

With anthropomorphization and possible deception, the review captures a closely related problem. The more human a system appears, the more easily people may attribute to it abilities, care, or confidentiality that do not follow from a fluent response. This does not mean that every human-like design deceives. In our view, however, it becomes ethically delicate as soon as the surface conceals institutional facts: that a provider operates the system, that algorithms and data flows are not neutral, and that the machine itself cannot vouch for any promise. Transparency must therefore do more than provide a one-time indication that one is speaking with AI.

The role must be established before the dialogue

Coghlan and colleagues situate ethical responsibility along the entire technology pipeline. Two years later, Meadi and colleagues show how broadly the questions arising from this are now being discussed and identify further research needs: risks and benefits compared to human professionals, appropriate roles in care contexts, effects on access, and the attribution of responsibility. Zhang and Wang complement the functional perspective by describing possible support services while at the same time highlighting limitations in long-term effects, memory, biases, genuine empathy, and ethical judgment.

From this follows our central position: An AI system for conversations about psychological distress must not be defined primarily by its linguistic persuasiveness. First, its institutional role must be limited. An information offering, a tool for professionals, a system for recording conversation trajectories, and a conversation service staged as a counterpart are not ethically the same. They generate different expectations and require different forms of oversight and accountability.

The most dangerous gap therefore does not necessarily lie in an awkward response. It arises when a system presents closeness without there being a recognizable responsibility behind that presentation. Software can speak around the clock. It cannot bear responsibility. Those who develop, offer, or integrate such conversations into professional structures must keep this difference visible, rather than blurring it through increasingly human-sounding language.

Sources & further reading