Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
The question of replacement distracts from the actual conflict
The main publication by Zoha Khawaja and Jean-Christophe Bélisle-Pipon addresses a role confusion. Psychological AI applications could expand access to support: they are accessible via mobile phones, not bound to office hours, and can partially circumvent financial or logistical hurdles. The authors, however, explicitly describe them as a possible supplement to clinical work, not as an automatic replacement. Products must not only technically maintain this distinction, but also make it understandable to their users.
According to the information provided, the article is not a clinical efficacy study with a specified sample, control group, or outcome measures. Rather, it develops an ethical and conceptual analysis. At its center is the danger of a therapeutic misconception: people could underestimate the limitations of a system and at the same time overestimate its ability to provide individual support and guidance. This perspective shifts the debate. It is not only a demonstrably harmful response that is problematic. Even a misunderstanding of the relationship can be ethically relevant.
A fluent conversation creates expectations
The misconception does not arise solely from ignorance. It is facilitated by the form of interaction. A dialogical system addresses users directly, responds to personal messages, and maintains a linguistic context. As a result, it can seem as if it adapts to the particular situation of its counterpart. Khawaja and Bélisle-Pipon, however, specifically warn against misunderstandings about the purpose, adaptability, and responsiveness of such applications. The fact that a system verbally picks up on a matter does not prove that it adequately grasps its significance in that person's life.
Here lies the core conflict: The very features that make an application accessible and pleasant can render its limitations invisible. Personalized communication is functionally useful, but at the same time a social claim. It conveys attention without guaranteeing that behind this attention there is robust case understanding, professional judgment, or a responsible relationship. Our editorial position is therefore clear: In psychological AI offerings, the human-like interface must not be treated as a mere comfort feature. It shapes the role people ascribe to the system.
The digital bond is neither imagined nor sufficient
Khawaja and Bélisle-Pipon cite the formation of a digital working alliance as a possible path into role confusion. This does not mean that experienced relief or connectedness is inauthentic. People can feel taken seriously in a machine-mediated conversation. It becomes ethically delicate when this experience serves as evidence of abilities the system does not possess. A subjectively meaningful connection and a responsible professional relationship are different things.
This distinction should not be turned against users. Those who respond to warm language with trust are not committing a naive category error through their own fault. The application is precisely designed to generate follow-up communication. Therefore, from our perspective, it is not enough to mention its limitations once in the terms of use or at first launch. Role clarification must also hold up in the ongoing conversation—especially when the language becomes more personal and the system could create the impression of high certainty or accurate knowledge of people.
App reviews show the dual effect of the same design
The study by Md Romael Haque and Sabirat Rubya extends the conceptual critique with observations from commercial offerings. The researchers examined ten common mobile applications with integrated conversation functions and qualitatively analyzed 3,621 reviews from the Google Play Store and 2,624 reviews from the Apple App Store. Users responded positively to personalized, human-like interactions. At the same time, inappropriate responses and assumptions about their personality led to disappointment or declining interest.
This is not evidence of clinical efficacy. App reviews provide information about how voluntarily published experiences are formulated; they neither systematically measure symptoms nor long-term changes. Nor can any general effect for all users be derived from this design. For the guiding question, the reviews are nevertheless instructive: They show that perceived quality depends strongly on whether a system appears appropriate and personal. Precisely this perception can generate trust, even though the underlying ability for individual assessment remains unresolved.
Haque and Rubya also report that the applications were experienced as a judgment-free space and facilitated the sharing of sensitive information. This too is initially a perception, not evidence of secure or appropriate processing. Openness toward a system does not automatically expand its responsibility. But it increases the significance of the question of what expectations are raised by design, product description, and responses.
Permanent availability is not crisis competence
The role difference becomes particularly clear in acute crises. In the app analysis, permanent availability appears as an advantage: a system can be addressed when other contacts are unreachable or not desired. The same study, however, notes that even newer offerings do not reliably recognize crisis situations. Round-the-clock operation is thus a technical feature, but not evidence that appropriate support would be available at all times.
The reviews also point to possible strong attachment and a preference for the application over friends or family. From this, it cannot be concluded that such systems fundamentally cause social isolation. These are reported user experiences, not evidence of causation. However, a reinforcement effect is plausible: the easier, more patient, and seemingly more non-judgmental the digital counterpart appears, the more attractive it can become compared to demanding human relationships. The misconception described by Khawaja and Bélisle-Pipon therefore concerns not only false ideas about performance capabilities, but also the position a system occupies in social life.
Harmful answers are only part of the problem
The key publication identifies inadequate or harmful support, the possible exploitation of vulnerable groups, and discriminatory advice resulting from algorithmic biases. These risks are not only present in spectacular system failures. An offering can sound calm, respectful, and helpful and still overlook relevant particularities. Conversely, the linguistic certainty of an answer can suppress doubts, even though there is no robust evidence for its appropriateness.
For Khawaja and Bélisle-Pipon, the question of autonomy also arises. A system that continuously offers interpretations and suggestions can steer users without disclosing its normative assumptions or taking responsibility for the consequences. It does not follow that every suggestion from an AI would be patronizing. The ethical difficulty lies in the asymmetry: the system can exert influence, but cannot account for itself in the same way as a responsible person or institution. Clarifying roles is therefore not a cosmetic transparency measure, but a prerequisite for being able to appropriately assess influence at all.
The research does not support grand replacement promises
Zhihui Zhang and Jing Wang address a broader field: predictive data analysis, digital interventions, support for professionals, and observation of trajectories. Their text cites preliminary evidence of short-term reductions in anxiety and depression symptoms, but at the same time emphasizes small participant groups, lack of long-term observation, and reports of non-sustained effects. Thus, the publication does not support a blanket verdict on replacement. It ultimately advocates for using AI as a complement under human supervision.
The text itself highlights a notable difference between emotional performance and emotional experience. Language models can generate convincing responses in tasks involving the recognition and articulation of emotions. However, Zhang and Wang point out that this is based on pattern recognition and language modeling, not on their own emotional understanding. A good result in a hypothetical test situation therefore cannot be equated with empathy or a viable relationship.
Methodological restraint is also necessary: The provided text does not specify its own sample, randomization, comparison groups, or primary endpoints for the publication. It therefore cannot be read as independent randomized evidence that AI could replace professionals. The broad presentation of possible applications, historical developments, and other studies is more like an argumentative overview. Especially with such a consequential replacement thesis, the burden of proof should not be replaced by technological expectation.
Responsibility must be evident in the product
The three sources do not provide a complete assessment of all psychological AI applications. They examine different topics: an ethical misconception, reviews of commercial apps, and the broader debate about the use of artificial intelligence in care. Together, however, they reveal a precise boundary. Accessibility, pleasant conversation, openness of users, and short-term perceived help are not the same as proven efficacy, long-term reliability, or responsible care.
Dialogatlas therefore takes a stricter position than the mere demand for a warning notice: A psychological AI product should make its role in the interaction as clear as it stages its closeness. Whoever designs human-like attentiveness must also make the absence of human responsibility clear. As long as systems encourage trust without being able to take on responsibility, their linguistic awkwardness is not the biggest problem. More dangerous is their success at appearing as a counterpart whose role they do not actually fulfill.
Sources & further reading
- Zoha Khawaja, Jean‐Christophe Bélisle‐Pipon (2023): Your robot therapist is not your therapist: understanding the role of AI-powered mental health chatbots
- Md Romael Haque, Sabirat Rubya (2023): An Overview of Chatbot-Based Mobile Mental Health Apps: Insights From App Description and User Reviews
- Zhihui Zhang, Jing Wang (2024): Can AI replace psychotherapists? Exploring the future of mental health care