Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
The obvious touchstone is too small
In dialogic AI, attention almost inevitably focuses on the individual response. Is it understandable, empathetic, and appropriate? Does the system maintain the thread of conversation? Such questions are legitimate, but they initially capture an interaction quality. They neither prove that symptoms decline, nor that a change persists, nor that the offering functions safely under everyday conditions. Especially in the field of mental health, linguistic persuasiveness must not be confused with proven efficacy.
The core conflict therefore lies between surface and infrastructure. Users experience a conversation; however, it is an entire arrangement of software, data processing, duration of use, responsibilities, and possible transitions to human help that becomes effective or harmful. A correlation between positive evaluation and continued use would not yet be evidence of efficacy. Likewise, high usage alone would not be proof of quality. Our editorial assessment is therefore deliberately strict: the conversation must be evaluated, but it must not be the unit to which all quality judgments are narrowed.
An overview of a field, not a judgment about a product
Source 1 is a review of digital mental health. It considers not only generative AI and large language models, but also smartphone apps, digital phenotyping, and virtual reality. Its subject is thus larger than the question of whether a single dialogue system gives good answers. The authors organize technological developments, the state of research on apps, problems of use, implementation issues, and inequalities in access into five interconnected themes.
This design also sets a knowledge boundary. The publication is not an examination of a specific AI offering, and the source evidence provided does not report a single sample, a uniform intervention, or a common effect size. Its value lies in the field perspective: it connects the scientific foundation of digital offerings with their real-world applicability. This does not yield a blanket efficacy judgment about generative AI. It does, however, yield a methodological warning: testing only a model’s outputs does not yet examine whether a viable form of mental health care emerges from them.
More studies are not enough if they examine the wrong situation
Torous and colleagues address the evidence on smartphone apps in several application areas, including well-being, depression, anxiety, schizophrenia, eating disorders, and substance use disorders. From this broad review, they derive the need for a new generation of more rigorous studies, including placebo-controlled and real-world studies. This is directly relevant to AI conversations: a comparison with no support at all answers a different question than a comparison with a digital but not therapeutically effective conversational offering.
Real-world relevance is more than a methodological addition. A system can appear plausible under controlled conditions and yet be rarely opened, abandoned early, or used outside intended workflows in everyday life. Conversely, a popular offering does not prove that it improves mental health. The source evidence does not provide a general ranking of individual products. It does, however, convincingly shift the burden of proof: clinical or care-related claims require appropriate comparison conditions, prospective testing, and results that go beyond liking, novelty, and perceived empathy.
Actual use begins after the first opening
According to the lead publication, a recurring problem with digital offerings is engagement. This does not merely mean whether a service is fundamentally accessible, but whether people continue to use it in a helpful way. As possible answers, the review mentions human support, digital navigators, context-dependent adaptive interventions, and personalized approaches. This reveals an uncomfortable insight: precisely those offerings that appear automatable may require additional human and organizational work to become practically relevant.
This work should not be understood as a subsequent repair of an otherwise finished product. It belongs to product and care design. Who assists with setup? Who helps interpret results? What happens when use is discontinued or the offering does not fit the situation? The sources do not answer these concrete responsibility questions for every system. They do show, however, that usage barriers are not solely characteristics of supposedly unmotivated users. Design, context, and lack of embedding also contribute.
Generative AI exacerbates the confusion between impression and effect
Source 2 specifically focuses on generative AI in mental health care. Torous and Blease distinguish foreseeable support functions such as documentation, administrative work, training, and routine symptom monitoring from more ambitious promises regarding prevention, diagnosis, and treatment. For text-based support, they refer to research with nonclinical samples in which perceived empathy was the focus, not clinical outcomes. This is not a trivial finding, but it is only an early step between technical feasibility, acceptance, efficacy, and actual effectiveness.
Language models make this distinction particularly difficult because correct and incorrect components can coexist in a fluent response. Source 2 illustrates this problem with research on treatment recommendations in oncology and warns against drawing conclusions about real diagnostic performance from standardized exam questions or simple case vignettes. The comparison does not replace an examination of mental health care, but it marks a structural risk: plausible language can make errors hard to detect. From an editorial perspective, evaluations must therefore not only measure average quality, but also examine which errors occur and whether people can notice them.
A standalone tool can also fragment care
The lead publication emphasizes the involvement of professionals, integration into existing services, and scalable delivery models. Source 2 formulates the same conflict more pointedly: isolated self-help offerings can fragment care, are not automatically scalable, and often remain unsustainable. The technical interface is therefore only one of several. Equally important are clinical workflows, interoperability, qualification, regulation, and the question of what role patients, relatives, administration, and development play in design.
It does not follow from this that every AI conversation must be directly connected to an institution. Such a general requirement is not supported by the three sources. What does follow, however, is that providers should not leave the intended context of use open. A tool for formulating thoughts poses different requirements than a system that observes symptoms or suggests diagnostic and treatment-related statements. The greater the claim, the less acceptable unclear responsibility. Our position: scaling does not mean merely providing the same interface to many people, but also scaling reliable transitions and responsibilities under real conditions.
Equitable access does not arise from mere availability
Source 1 explicitly addresses the question of whether digital innovations work for different population groups. It refers to the adaptation of tools for historically marginalized groups as well as for low- and middle-income countries. Co-design and implementation science appear here not as decorative buzzwords, but as ways to bring together scientific quality and practical usability. A globally accessible language model is therefore not yet a globally suitable offering.
Source 2 supplements this perspective with biases, privacy, and the origin of training data. The authors warn against stigmatizing depictions of mental illnesses and against the handling of sensitive health information by general systems. The state of the sources does not allow the claim that all systems show the same errors to the same extent. The fundamental consequence is nevertheless clear: cultural adaptation, data protection, and fair performance must not be examined only after an offering has already been distributed. They help determine who can express themselves, whose language is understood, and who bears risks.
The longer story protects against the novelty effect
The 2021 review by Boucher and colleagues shows that AI-powered dialogue systems were already integrated into digital interventions before the current boom. These include diagnostics and screening, symptom management and behavior change, as well as content delivery. The authors highlight research needs regarding perceptions of AI, individual differences, privacy, and ethics. Their outlook on dynamic, highly personalized systems seems relevant today, but in the present abstract it provides no evidence that such a future has already been effectively or safely realized.
Read together, the three publications shift the object of assessment. The individual response does not disappear, but it becomes part of a longer chain: data and model, interaction, use over time, human support, transitions, organization, and distribution of benefits and risks. This is precisely where Dialogatlas’s decisive editorial thesis lies. The best AI conversation is not automatically the most human or the most eloquent. In the context of mental health, it is the one whose purpose is limited, whose evidence matches its claim, and whose place in the broader process is comprehensible. The visible conversation should therefore never obscure the fact that the actual design task continues beyond the last sentence.
Sources & further reading
- John Torous, Jake Linardon, Simon B. Goldberg et al. (2025): The evolving field of digital mental health: current evidence and implementation issues for smartphone apps, generative artificial intelligence, and virtual reality
- John Torous, Charlotte Blease (2024): Generative artificial intelligence in mental health care: potential benefits and current challenges
- Eliane M. Boucher, Nicole Harake, Haley Ward (2021): Artificially intelligent chatbots in digital mental health interventions: a review