Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
15 studies map a heterogeneous field
Casu and colleagues searched several databases, including MEDLINE, Scopus, and PsycNet. Additionally, they used AI-assisted search tools such as Microsoft Copilot and Consensus. Studies were selected based on predefined criteria; several researchers independently performed data extraction and quality assessment. In the end, 15 studies met the requirements of the scoping review.
The fields of application were broad. The studies covered support during the COVID-19 pandemic, offerings for depression, anxiety, and substance use, preventive approaches, health promotion, and usability. The review describes possible improvements in mental and emotional well-being, contributions regarding specific burdens, and support for behavior change.
A scoping review primarily maps what was investigated and which topics become visible. It does not calculate a uniform overall effect for all applications. With such diverse target groups, systems, and endpoints, a blanket judgment would be misleading anyway. The value of the work lies precisely in naming, alongside possible effects, the conditions of actual use.
Efficacy and feasibility are different questions
A system can show a positive effect under study conditions and still be barely used in everyday life. Conversely, a chatbot can be pleasant and easy to use without bringing about the intended change. Efficacy asks about results. Feasibility asks whether the offering can be accepted, understood, operated, and meaningfully embedded under real-world conditions.
This distinction seems self-evident, but is often lost in public debate. Good reviews of the interface are read as efficacy; small symptom changes are presented as proof of a viable product. The lead publication brings both levels together. It sees potential and at the same time names usage, engagement, and integration into care as open challenges.
For development, this means: a convincing model comparison solves only part of the problem. People must recognize what a system is intended for without being put off by long explanations. They must be able to start a conversation, want to return, and understand what happens to their data and previous conversations.
Engagement is not a decorative metric
Engagement is often presented in studies and product reports as duration of use, number of sessions, or completion rate. These metrics say something about whether an offering is used, but do not automatically explain why. People may stay for a long time out of interest, repeatedly start over out of frustration, or experience a short conversation as entirely sufficient.
In mental-health conversations, a qualitative dimension is added. A system can lose engagement because it gives advice too early, repeats itself, forgets corrections, or turns every sentence into a problem. Such errors are not merely matters of style. They change whether a person feels taken seriously in the conversation and whether they want to express another thought.
Dialogatlas therefore considers a mere increase in duration of use to be a poor universal goal. A good conversation may be short. What is more relevant is whether the person receives the chosen type of conversation without being manipulated into staying. Engagement should be understood as voluntary fit, not as the longest possible retention.
The meta-analysis shows benefit with a clear limit
The systematic review and meta-analysis by Han Li and colleagues complements the scoping review quantitatively. It searched twelve databases and identified 35 suitable studies from 7,834 records for the systematic review. Fifteen randomized controlled trials were included in the meta-analysis.
The pooled results showed significant reductions in depressive symptoms and distress. For general psychological well-being, however, no significant improvement was found. This difference is important. A conversational agent can have an effect on certain measures of burden without thereby producing comprehensive well-being.
In the analysis, the effects were stronger, among other things, for multimodal and generative systems, when embedded in mobile or instant messaging applications, and in certain target groups. Such moderator findings are interesting indications. However, they are not a ready-made blueprint, because the number of studies and the differences between the applications must be taken into account.
Personalization needs more than a name
Casu and colleagues identify personalization and context-specific adaptation as central development areas. In many products, personalization is implemented superficially: the system knows the first name, remembers a selected goal, or uses a preferred tone. That can be pleasant, but it does not yet get to the core.
Context also means knowing what kind of conversation is currently desired. Does someone want to tell their story first, better understand a situation, hear a different perspective, or look for a next step? A system that ignores these differences can give professionally plausible answers and yet constantly talk past the user's actual needs.
Real personalization must also not become secret psychological profiling. It needs comprehensible settings, correctable memories, and a clear limitation of what is stored and interpreted. More data does not automatically mean more understanding.
Integration means more than just an interface
The lead publication describes integration into existing healthcare systems as a challenge. Technically, this could be due to data formats, responsibilities, or workflows. Substantively, the question is larger: What role should the chatbot take on in relation to other offerings?
A freely accessible conversational partner, a structured self-help program, a tool in a treatment, and a clinical monitoring system require different rules. They must not be grouped together solely because they all use a chat interface. The more closely a product is integrated into care, the more important defined handovers, responsibilities, and application-specific evidence become.
At the same time, there are meaningful applications outside a clinical pathway. Not every conversation that provides relief requires institutional referral. Honesty about the role is crucial. Integration can also mean shaping one's own boundaries so clearly that a low-threshold offering does not inadvertently appear as treatment.
Human-AI integration is not automatically a hybrid model
For future research, the scoping review calls for larger studies and clarification of optimal human-AI integration. This term is quickly equated with a hybrid model in which professionals take over every critical situation. That can be useful in certain applications, but it is organizationally and economically demanding.
Human involvement can also start earlier: in defining the task, reviewing conversation errors, selecting test cases, and deciding which data are processed at all. Not every conversation needs to be monitored live for humans to take responsibility for the system.
The second comparison source discusses similar possibilities and limitations on a broader level. It emphasizes accessibility and scalability, but mentions limited long-term evidence, bias, and human oversight. The decisive point is not to put a human behind the AI everywhere. It is to not hand over responsibility to the model's surface.
The best answer needs a good home
The 15 studies in the lead publication and the meta-analysis with 35 included works justify more than the claim that mental chatbots are merely a technical trend. There is evidence of benefit for certain symptoms and burdens. At the same time, long-term effects, general well-being, usage, and safe integration remain open.
Our editorial conclusion is practical: The quality of a mental chatbot does not arise solely from the model. It arises from the task, conversation design, memory, data protection, interface, error monitoring, and a realistic role. A strong model can facilitate this work, but not replace it.
Therefore, such a system rarely fails solely because of its first response. It fails because the system forgets the chosen conversation mode in its next response, because settings have no effect, because no one learns from recurring errors, or because the product promises more than its studies show. Conversely, that is precisely where the opportunity lies: Good product work can turn a possible response into a more reliable conversation.
Sources & further reading
- Mirko Casu, Sergio Triscari, Sebastiano Battiato (2024): AI Chatbots for Mental Health: A Scoping Review of Effectiveness, Feasibility, and Applications
- Zhihui Zhang, Jing Wang (2024): Can AI replace psychotherapists? Exploring the future of mental health care
- Han Li, Renwen Zhang, Yi‐Chieh Lee et al. (2023): Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being