Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
A Big Question with Too Little Resolution
Zhang and Wang explicitly place the replacement question at the center of their contribution. They cite potential advantages such as constant availability, scalability, consistent procedures, and easier access. At the same time, they describe limitations in long-term memory, algorithmic bias, data protection, ethical judgment, and genuine empathy. Their own outlook therefore does not point to complete substitution but to a supportive role under human supervision. The present source provides no direct randomized comparison between an AI and psychotherapists that could support a claim of replacement.
The problem lies in the category itself. A profession does not consist of a single task. It encompasses conversation management, assessment, documentation, progress monitoring, decisions under uncertainty, and dealing with changing situations. A machine can take over or support individual tasks without thereby filling an entire professional role. From our perspective, the binary question of replacement or non-replacement therefore obscures the truly testable questions: What specific function should the system fulfill, with what data does it work, and how is its success measured?
85 Studies, Three Fundamentally Different Application Areas
The lead publication by Cruz-Gonzalez and colleagues offers the strongest basis for this refinement. The team searched CCTR, CINAHL, PsycINFO, PubMed, and Scopus from the beginning of each database’s holdings to February 2024. According to predefined inclusion criteria, 85 studies were included in the systematic review. Applications from the areas of diagnosis, monitoring, and intervention were examined. This design is important: the review does not simply assess whether AI works in mental health care, but rather assigns different technical methods to different tasks.
The methods commonly used also differed by area of application. For diagnostic tasks, support vector machines and random forests dominated. For monitoring, machine learning was the most common category; for interventions, it was AI-based conversational systems. This finding alone contradicts the popular image that mental-health AI consists mainly of a talking interface. A large part of the field classifies data, estimates risks, predicts reactions, or tracks trajectories. The conversation is visible and easy to demonstrate; the systematic review, however, reveals a broader technical structure.
Accuracy does not yet answer a question of care
According to the review, AI tools appeared suitable for detecting and classifying mental disorders, predicting risks, estimating responses to treatments, and monitoring ongoing trajectories. This is a substantially positive finding and should not be downplayed out of mere technological skepticism. The provided information, however, does not list common performance measures, pooled effect sizes, or a breakdown by disorder, data type, or application context. From the overall judgment, it therefore cannot be inferred that every included method is clinically reliable or transferable to other populations.
Above all, the target measures are not interchangeable. A good classification initially shows that a model can distinguish categories in a specific dataset. An accurate risk prediction, in turn, is different from a demonstrated improvement in the subsequent course. And a convincingly formulated response proves neither diagnostic accuracy nor a lasting effect. Our assessment is therefore neither dismissive nor euphoric: the review documents serious technical capability, but it does not provide a blanket license for care-related claims.
Conversational systems change the evaluation task
In diagnostic classification, model output and reference category can at least in principle be compared directly. A dialogue-based intervention is harder to capture. It does not merely produce a prediction but immediately formulates content to which a person can react. In doing so, comprehensibility, emotional appropriateness, symptom change, and temporal stability can diverge. Cruz-Gonzalez and colleagues identify conversational systems as the most common AI method in the intervention area; the available information, however, does not state how many studies concern individual application types or how durable their results were.
It is precisely here that the talking interface becomes seductive. Linguistic fluency makes a performance immediately tangible, while errors in classification or prediction often remain invisible to users. From an editorial standpoint, we therefore consider it wrong to treat conversation quality as a proxy for efficacy. A system can sound attentive and still deliver unsuitable content. Conversely, a limited, not very human-like tool can reliably support a clearly defined task. Human-likeness is a product feature; it is not a uniform seal of quality.
Older language research explains the data dependence
The systematic review by Le Glaz and colleagues from 2020 sharpens this point. After a search in four medical databases, 58 of 327 identified articles were qualitatively analyzed. The studies used machine learning and natural language processing, among other things, to extract symptoms, classify severity levels, compare treatment outcomes, and obtain psychopathological indications. Medical records and social media were the two most important data sources; the populations studied included people in medical databases, individuals in emergency departments, and users of social media.
The authors saw this as access to previously underused information, such as everyday habits that are usually not accessible to clinicians. At the same time, they observed that powerful classifiers were preferred over models that function transparently. Many procedures tended to confirm existing clinical hypotheses rather than generate entirely new insights. Social media in particular constituted an imprecise cohort, and language-specific features made transfer to other languages difficult. The fact that the more recent review again demands more robust, more diverse datasets as well as more transparency and interpretability shows that these methodological questions have not disappeared with more powerful models.
The discernible benefit deserves precise language
Both systematic reviews yield a constructive finding. AI can identify patterns in extensive, previously underutilized data sets and provide useful information for detection, prognosis, or course observation. In the intervention area, the always-available communication also opens a plausible pathway to support. Zhang and Wang also cite anonymity and the interaction perceived by some people as less judgmental as possible reasons for more open statements. Such potentials are more than mere science fiction, even if their scope depends on the application and the state of the evidence.
When it comes to demonstrating efficacy, wording remains crucial. Zhang and Wang report preliminary studies on short-term improvements in anxiety and depression symptoms, but at the same time point to small groups, missing long-term follow-up, and findings that effects did not persist over longer periods. It does not follow from this that the interventions are ineffective or equivalent to human treatment. The appropriate conclusion is narrower: certain digital interventions can be helpful in the short term; whether this benefit remains stable and under which conditions it arises is not conclusively clarified with the information available here.
Emotional language is a skill, not an inner life
The contribution by Zhang and Wang illustrates another conceptual trap. It refers to studies in which ChatGPT generated responses at or above the level of the general population on tasks of emotional perception. At the same time, the authors make clear that this performance is based on pattern recognition and language modeling, not on experienced emotions. The ability is by no means worthless: a system that linguistically recognizes emotional components and responds to them appropriately can be more useful in certain situations than one that cannot.
But it does not prove the sameness of the relationship. Zhang and Wang contrast the linguistic abilities with missing genuine empathy, difficulties with long-term continuity, and limits in capturing nonverbal or culturally shaped nuances. This gives rise to a methodological conflict for assessment: a test can measure the quality of emotionally formulated responses, but not automatically the viability of a longer-term professional relationship. Whoever infers from a good language benchmark that a profession can be replaced imperceptibly shifts the level of evidence.
The reliable unit is the limited task
Cruz-Gonzalez and colleagues recommend more diverse and robust datasets as well as more transparent and more interpretable models. Combined with the two comparison sources, this does not amount to a general rejection but rather a more precise research logic. Diagnosis, monitoring, and intervention each require their own target variables. Even within these areas, it must remain clear for which population, language, data source, and time period a result applies. The search period of the main review ends in February 2024; statements about systems published later or announced model generations cannot be derived from it.
Our editorial position is therefore clear: The future of mental health AI is not determined by whether a system overall comes across as a professional. What matters is which limited task it demonstrably fulfills and which claim is actually supported by the respective study design. A model can classify risks well without improving the course of the condition. A dialogue can provide short-term relief without ensuring long-term continuity. These differences do not diminish the potential. They make it describable at all – and prevent an impressive conversation interface from being declared the answer to questions that were never tested.
Sources & further reading
- Pablo Cruz-Gonzalez, Anxun He, Eva K. M. Lam et al. (2025): Artificial intelligence in mental health care: a systematic review of diagnosis, monitoring, and intervention applications
- Aziliz Le Glaz, Yannis Haralambous, Deok-Hee Kim-Dufor et al. (2020): Machine Learning and Natural Language Processing in Mental Health: Systematic Review
- Zhihui Zhang, Jing Wang (2024): Can AI replace psychotherapists? Exploring the future of mental health care