Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The experience is a finding – but of what kind?

Siddals, Torous, and Coxon examined experiences from real-world use of generative AI systems for psychological concerns. To this end, they interviewed 19 people. The focus was therefore not on prescribed exercises under controlled conditions, but on the meanings that people attributed to their own conversations with generative AI. This is particularly relevant for product development: The study shows what available systems are actually used for and what qualities users recognize in them.

Such an interview study can show in detail how support is experienced and described. However, it cannot determine whether a system reliably reduces symptoms, lowers risks compared with other offerings, or has caused changes. Reports of positive outcomes are neither trivial nor objective evidence of efficacy. The key editorial distinction is therefore: Subjective significance is a research topic in its own right, but not a substitute for clinical testing.

Four experiences explain the strong attachment

The researchers developed four themes from the interviews. First, participants described an emotional refuge: The conversation offered a space in which they could open up. Second, they experienced the responses as insightful guidance, especially in relationships. Third, they reported joy in the connection itself. Fourth, they explicitly compared the system with human psychotherapeutic support. Thus, the AI does not appear merely as an information tool. It is placed in a social and sometimes care-adjacent role.

Participants also reported positive effects, including better relationships and experiences they associated with processing trauma or loss. The wording is important: what is documented is their accounts, not independently established changes. Neither standardized outcome measurements nor a control group nor a long-term comparison are evident from the provided abstract. Nor is it proven whether the same people would have found other support without the system. The value of the study therefore lies in the precise description of a phenomenon, not in evidence of efficacy.

Why the conversation can function as a protected space

The reported openness is plausible. A counterpart that is always reachable requires no appointment, shows no visible irritation, and responds immediately. Generative systems can also respond to free-form accounts instead of only guiding users through fixed menus. This can create the impression of actually being heard. The interviews suggest that precisely this combination of accessibility, linguistic adaptation, and the absence of social sanction gave the conversations meaning.

But the impression of a protected space and an actually protected space are not the same. Pleasant communication does not guarantee confidentiality, professional competence, or reliable crisis response. Nor does an empathetic-sounding answer prove that the system understands a person’s situation. Our assessment is therefore: The social effect of language must be treated as a real product property, even if there is no sentient counterpart behind it. Those who look only at technical functionality underestimate the power of attachment; those who equate attachment with human understanding overestimate the system.

Users themselves identify the unfinished role

It is noteworthy that the respondents did not only describe closeness and usefulness. They also demanded better safety precautions, a more human-like memory, and the system’s ability to guide the process more strongly. These wishes point to an internal contradiction. The system is valued precisely for its openness and flexibility, but is also expected to offer continuity, judgment, and direction – qualities to which considerable responsibility is attached.

More memory, for example, would not automatically be just an improvement. It could make conversations more coherent, but would inevitably affect the handling of particularly sensitive information. More guidance could provide orientation, but would shift the AI further from a reactive writing tool to an intervening authority. The lead study does not prove how these conflicts are to be resolved. Rather, it shows that usage demands push the systems into a role for which fluent language alone does not constitute a sufficient basis.

Relationship does not arise from a single design language

The diary study by Xu, Lee, Stasiak, and colleagues extends this finding with a temporal perspective. 26 adults used the Woebot and Wysa offerings over four weeks. Data consisted of weekly surveys, conversation screenshots, and semi-structured interviews; analysis was conducted using reflexive thematic analysis. 18 participants reported a pronounced or weaker bond with at least one of the two systems. Others did not develop such a relationship.

What mattered was evidently not a universally optimal form of conversation, but rather the fit. The analysis cites, among other things, the desire to lead a conversation oneself or to be led, the alignment between preferred modes of expression and possible inputs, expectations of caring, the perceived usefulness of advice and activities, colloquial communication, and private, non-judgmental conversations. People with both lower and higher scores on the WHO-5 well-being index reported bonds. It does not follow from this that psychological well-being is meaningless; the study merely suggests that the capacity for bonding was not strictly tied to it.

Fit is a design goal, not a measure of efficacy

Taken together, the two qualitative studies argue against a simple formula such as: the more human the system, the better the relationship. Some people want to lead, others want to be led. Some expect warmth, others primarily relevant guidance. A product can therefore feel engaging and helpful to one person, while the same conversational logic comes across as mechanical, patronizing, or unsuitable to another.

For research and development, this is a methodological note. Bonding, satisfaction, repeated use, and perceived usefulness should not be merged into a single concept of success. A strong bond can stabilize usage, but it neither proves a health improvement nor appropriate decisions by the system. Conversely, a utilitarian tool does not need to appear relationship-like to serve a limited purpose. Our position is: the intensity of the relationship must not become a surrogate indicator for quality.

The replacement question promises more clarity than the research yields

Zhang and Wang explicitly raise the question of whether AI could replace psychotherapists. Their contribution assesses possible functions: better accessibility and scaling, continuous availability, support in evaluation and personalization, and potentially more open self-disclosure to machines. At the same time, they address the lack of genuine empathy, limited long-term memory, algorithmic biases, data protection, autonomy, and the need for human oversight. They also point out that previous studies often involve small groups and no long-term follow-up, and that short-term effects do not readily persist.

The provided source excerpt, however, does not report its own sample, randomization, allocation, or outcome analysis. Despite the accompanying assessment as a randomized study, no corresponding study design can be reconstructed from this material. The text is therefore to be read here as a broadly argued literature contribution, not as direct evidence of replacement. Moreover, abilities on linguistic emotion-recognition tasks do not mean that a model experiences emotions or understands a person in their life situation. The big replacement question thus obscures the more precise task: For which narrowly defined functions is there robust evidence under which conditions?

The product promise must remain smaller than the feeling

The three publications do not yield a verdict on generative AI as an entire care category. The systems, usage patterns, and study designs differ too greatly for that. The lead study provides solid evidence that individual people experience such conversations as exceptionally fitting and consequential. The diary study additionally shows how much attachment depends on expectations and interaction preferences. The contribution by Zhang and Wang makes clear how many technical, ethical, and clinical questions stand in the way of a claim of substitution.

Dialogatlas draws a deliberately restrained conclusion from this: The more a product stages emotional refuge, continuity, and guidance, the less it may fall back on the note that it merely generates text. Its design shapes expectations and can give rise to a care-like relationship. Precisely for this reason, providers should not claim more certainty, understanding, or effect than has been tested. To downplay users' experiences would be wrong. To pass them off as evidence of reliable mental health care would be equally wrong. Responsible development begins in this difference.

Sources & further reading