Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The role name changes the standard

The review’s guiding question is narrower than the broad term AI in mental health care suggests. It examines ethical considerations regarding conversational AI that appears to people with mental health problems in the role of a therapist. This means systems that interact with a person and formulate their outputs using artificial intelligence. Thus, it is not equally about prognostic models, administrative software, or digital questionnaires, but about linguistic systems to which a particularly demanding function is attributed.

That is precisely where the core conflict lies. A tool for conveying content must be assessed differently than a system that builds trust, receives crisis statements, or creates the impression of professional competence. The more a product is staged as an independent counterpart, the less the statement that it is merely software suffices. From an editorial perspective, the role designation must therefore not be left to marketing or the spontaneous perception of users. It determines which expectations are legitimate and who must answer for borderline cases.

A map of the discourse, not an efficacy judgment

Meadi and colleagues systematically searched PubMed, Embase, APA PsycINFO, Web of Science, Scopus, the Philosopher’s Index, and the ACM Digital Library. The search combined embodied artificial intelligence, ethics, and mental health; further texts were added via snowball search. Publications in English or Dutch of various types were considered, excluding symposium abstracts. Two researchers independently assessed eligibility. For the analysis, an initially expectation-based extraction scheme was revised and supplemented during the charting process.

101 publications were included, 96 of them from 2018 or later. Reviews formed the largest single group with 22 contributions, followed by 17 commentaries. This is crucial for interpretation: The review maps which problems occur in a young field of debate. It does not measure how frequently certain harms actually occur in care, nor does it compare treatment outcomes. In addition, there is a limitation noted by the authors: The perspectives of relevant stakeholders are insufficiently represented to date. The map is extensive, but it should not be confused with the terrain.

Ten topics are not a ranking of real dangers

A concern was designated as its own topic if it appeared in more than two publications. This resulted in ten areas. The most frequently addressed were privacy and confidentiality with 62 of 101 texts, safety and harm with 52, justice with 41, and efficacy with 38. These are followed by responsibility and accountability with 31, empathy and humanity with 29, explainability, transparency, and trust with 26, and anthropomorphization and deception with 24. Concerns about jobs in healthcare appeared in 16 contributions, and autonomy in 12.

These numbers describe attention in the literature, not the magnitude of a risk. A topic can be discussed often because it is visible early or conceptually well established. A less frequently addressed problem can still be consequential. The comparatively low presence of the autonomy question should therefore not be reassuring. The strength of the review lies in the joint presentation of the conflicts: Confidentiality, safety, or justice are not isolated checkpoints, but interlock as soon as a system engages in sensitive conversations and actions or expectations emerge from them.

In a crisis, communication becomes responsibility

Under safety and harm, the lead publication particularly names suicidality and crisis management, harmful or incorrect suggestions, and the risk of dependence on conversational AI. The review thus documents that these dangers are discussed in the literature; it provides no event rate and no evidence that all examined systems are equally risky. Nevertheless, the combination of topics is revealing. An incorrect answer becomes more serious when users consider the system to be competent, caring, or reliable.

Here, safety directly touches the question of responsibility. Who is accountable when the system misjudges a critical statement, reacts inappropriately, or conveys the impression that it has taken over the situation? The review identifies responsibility and accountability as a separate topic and calls for their clarification as a research need. Our assessment goes one step further: An application must not linguistically simulate comprehensive responsibility if organizationally no one bears that responsibility. A warning note in the margin does not automatically resolve the contradiction between the staging of a relationship and the actual responsibility structure.

Closeness is both surface and data relationship

Privacy and confidentiality are the most discussed area of the review. That is not surprising: a conversation can only feel personal if people bring in personal information. At the same time, the literature addresses explainability, transparency, and trust, including the impact of opaque black-box algorithms. Trust therefore concerns not only the tone of a response. It also concerns whether users can understand what kind of system they are talking to and on what its reactions are based.

Empathy and humanity, as well as anthropomorphization and deception, mark the other side of this closeness. The publication by Zhang and Wang emphasizes that a language model can generate emotional responses based on patterns without experiencing feelings. This is not merely a philosophical subtlety. Perceived empathy can facilitate use, but it is not evidence of emotional understanding, professional judgment, or a viable relationship. From an editorial perspective, we therefore consider a clear distinction indispensable: a convincing linguistic expression of compassion is a system property; its appropriate effect in care remains an open empirical and normative question.

Greater reach can create new inequality

The prospect of more easily accessible support is among the strongest arguments for conversational AI. Boucher and colleagues already described several fields of application for digital interventions in 2021: diagnostics and screening, symptom management, behavior change, and content delivery. Their overview discusses the potential of usable and accepted applications, but also points to open questions regarding perceptions of AI, individual differences, privacy, and ethics. A possible technical reach therefore does not yet imply equitable access.

In the scoping review, equity is one of the most frequent topics, with 41 contributions. Mentioned are, among others, health inequalities due to differing digital literacy. A service available around the clock can lower barriers while simultaneously disadvantaging those who understand it less well, trust it inappropriately strongly, or lack suitable access. Whether conversational AI reduces gaps in care cannot be answered from the three publications. The lead publication explicitly cites the effects on access as further research need. Accessibility is initially a product property; more equitable care would be an outcome that needs to be demonstrated.

The substitution debate is running ahead of the evidence

Efficacy appears in 38 of the 101 texts captured by the scoping review. This frequency, however, proves neither a benefit nor its duration. Boucher and colleagues speak of promising applications and outline future developments, but in the provided abstract they do not supply summary effect sizes, samples, or long-term results. Zhang and Wang point to preliminary indications of short-term reduction in anxiety and depression symptoms, while emphasizing small study groups, lack of long-term follow-up, and the need for large randomized controlled trials. They also note that short-term benefits do not necessarily persist over longer periods.

For source 3, there is an additional methodological limitation. It is described in the provided metadata as a randomized study; however, the supplied text does not document its own randomization, sample, control group, or independent outcome analysis. Instead, it unfolds a broad, partly speculative discussion of whether AI could take over psychotherapeutic functions. Therefore, on the basis of the available material, this publication cannot be treated as direct randomized evidence of substitution. Its cautious concluding position is nevertheless relevant: AI should complement rather than replace human elements and requires oversight as well as further clinical validation. The public-facing comparison of human versus machine is thus larger than the substantiated evidence base.

A limited function is more credible than an artificial counterpart

The review does not end with a finished ethics guideline. It identifies open tasks: examining risks and benefits compared with human professionals, determining appropriate roles in care contexts, assessing consequences for access, and clarifying accountability. This sequence should be taken seriously. The question is not, in the abstract, whether AI can conduct conversations. Rather, it is which concrete function is being evaluated under which expectations. Content delivery, screening, accompaniment, and independent therapeutic responsibility are not interchangeable product variants.

Our editorial position is therefore deliberately restrictive: conversational AI does not gain legitimacy by appearing as human-like as possible, but rather by having its role described more narrowly and honestly than its language might suggest. Only such a role framework makes measuring efficacy meaningful, because then it is clear what success, failure, and comparison are measured against. It also makes ethical conflicts more manageable without downplaying them. The most important achievement of the scoping review is not to enumerate ten known concerns. It shows that almost every one of them escalates as soon as a system signals relationship, while responsibility, evidence, and institutional integration remain undefined.

Sources & further reading