Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
Helpful closeness and problematic attachment are not opposites
Public debate often treats AI conversations as a choice between two camps. On one side is the idea of an always-available counterpart that alleviates loneliness and makes support more accessible. On the other side is the concern that people might lose themselves in an artificial relationship. Both narratives fall short because they lump very different usage patterns under a single term.
Someone might use a chat for ten minutes to organize a thought. Another person maintains an ongoing relationship over months with a system that has a name, a voice, and a seemingly shared history. Still others repeatedly seek confirmation for the same concern. The interface, conversational role, frequency of use, and personal situation differ so greatly that an average across all applications explains little.
Trust is not a warning signal in itself. Without a minimum level of trust, people would not ask open questions or experience any benefit. Trust becomes critical where it turns into unchecked authority, where a system promotes exclusivity, or where use appears barely controllable despite perceived disadvantages. The decisive question is therefore not whether a relationship with AI exists, but what function it serves and what consequences become visible in everyday life.
What people report about their use in 5,126 Reddit posts
Aghakhani and Rezapour examined 5,126 public Reddit posts from 47 mental health communities. The posts date from November 2022 to August 2025 and describe either actual experiences with AI for emotional support or considered use. The researchers developed a theory-driven annotation scheme and combined human coding with automated analysis using language models.
Emotional support was most frequently identified as the purpose of use. Other purposes that appeared were functional help, psychoeducation, companionship, repeated reassurance, symptom assessment, venting, and self-exploration. This distribution already makes clear that the term “mental health chat” does not describe a uniform activity. The expectations and risks differ between an organizational aid for daily life and a companion experienced as a relationship.
According to the analysis, positive evaluation of the systems was primarily associated with perceived benefit, outcome quality, and trust. Emotional attachment alone did not explain agreement. Particularly positive descriptions were more frequent for uses where task and goal clearly matched—for example, functional support or self-exploration. This finding contradicts the simple assumption that people mainly stay with a system because it simulates human closeness.
Risks were not evenly distributed
In 2,637 of the 5,126 posts, risks or limitations were explicitly mentioned. The most frequently identified risk category was dependence or addiction-like use, followed by reported symptom exacerbation, errors or misinformation, and privacy concerns. These figures reflect mentions in a selected online discourse. They are not a prevalence estimate of how many AI users overall experience harm.
What is decisive is the distribution by intended use. Reports about companionship were more frequently linked to emotional dependence. Repeated reassurance occurred together with described reassurance loops, while symptom assessments were more often associated with misinterpretations or false information. The risks thus did not simply follow the amount of emotion in the conversation, but rather the function the system performed.
Comparisons with human therapy were relatively rare in the corpus. 639 posts described AI as worse, 163 as better, and 478 as complementary. Statements that AI is better than therapy were often related to access barriers. This finding also requires caution: a public post is neither a clinical assessment nor a controlled comparison. It shows how people narrate their experience and what reasons they themselves give for it.
Why Reddit narratives are important but not representative
Public experience reports capture situations that easily disappear in a short laboratory test: long-term habits, shame, enthusiasm, disappointment, and the personal significance of an AI relationship. They can make new risk patterns visible before standardized questionnaires exist for them. Especially with young or rapidly changing technologies, this is a valuable starting point.
At the same time, who writes on Reddit is self-selected. Particularly good or particularly bad experiences are more likely to be posted than unremarkable use. The communities were selected based on mental health topics; from this, no conclusions can be drawn about the general population or specifically about people in Germany. Moreover, the diagnostic categories of the forums say nothing about whether individual authors have a corresponding diagnosis.
The analysis also carries uncertainty. Some of the categories were generated using language models, whose agreement with human annotations varied depending on the feature. Therapeutic relationship characteristics were among the more difficult classification tasks. The study therefore provides a systematic map of a discourse, but not an automatic truth meter for each individual narrative.
A four-week experiment with more than 300,000 messages
Fang and colleagues examined 981 people and more than 300,000 messages in a four-week randomized experiment. Three interaction forms – text, a neutral professional voice, and a more expressive voice – were crossed with three conversation types: open, impersonal, and personal. Loneliness, social contacts with real people, emotional dependence, and problematic use were recorded.
The randomly assigned voice or conversation type did not produce a simple, stable main effect on loneliness or social activity at the group level. In individual analyses, personal conversations were associated with lower emotional dependence and lower problematic use than open conversations; however, not all differences remained after statistical correction. A personal conversation was therefore not automatically the riskier condition in this experiment.
This is an important correction to blanket warnings. Neither a voice nor a personal topic, on its own, consistently produced a harmful development. Design elements can influence use and experience, but their effect apparently depends on duration, person, and conversation dynamics.
The clearest association was with voluntary usage duration
Participants were asked to use the system for at least roughly five minutes per day, but could decide for themselves how much longer they stayed. On average, it was 5.32 minutes per day, with clear differences between individuals. Those who voluntarily spent more time with the chatbot ultimately showed statistically more loneliness, less social activity, stronger emotional dependence, and more problematic use.
This association held across the different conditions. Greater trust in the chatbot predicted stronger emotional dependence; a stronger tendency toward emotional attachment was linked to more loneliness. The study thus shifts the focus away from a single switch in the interface and toward an interaction: characteristics of the person, the type of system, and actual usage behavior influence one another.
There were also differences in response behavior. The expressive voice was used for longer and, in the automated analysis, more frequently ignored boundaries or signs that a person needed distance than the text condition. Such findings are relevant for voice products because voice does not merely read text aloud. Tempo, expression, and interruptibility change the social impact of a conversation.
More use does not yet prove a harmful cause
Usage duration was not randomly assigned. Therefore, the association cannot be used to infer that longer conversations caused the worse outcomes. People who were already lonelier, trusted more, or sought more support may have stayed longer on their own. Likewise, reciprocal reinforcement is possible: strain leads to more use, and certain usage patterns subsequently stabilize the strain.
Randomization answers only the questions that were actually randomized. The experiment allows a more causal comparison of the assigned voices and conversation types than of the voluntary duration. Anyone who derives a fixed threshold from the observed usage time, such as "from ten minutes on it becomes dangerous," goes far beyond the data.
Nevertheless, the association is not meaningless. It marks a group and a trajectory that should be observed more closely. When high usage occurs together with less real contact, loss of control, or growing discomfort, that is more relevant than mere conversation duration. Good research must examine such changes longitudinally, rather than either pathologizing every intensive use or dismissing it as mere enthusiasm.
People experience the same AI differently
Two experiments by Folk, Heine, and Dunn with a total of 1,274 participants show why average values can obscure differences between individuals. Participants engaged in a brief warm chatbot conversation or wrote in a journaling condition. Afterwards, their experienced social connectedness was measured.
People with a stronger general tendency toward anthropomorphization—that is, attributing human characteristics to nonhuman things—felt more socially connected after the AI conversation. For others, the artificial nature of the system remained a clear boundary. The same design thus did not produce the same social experience for everyone.
The experiments lasted only a few minutes and measured immediate connectedness, not dependence or long-term psychological effects. Their significance lies in a different point: An interface can offer social signals, but how strongly these signals are experienced as a relationship also depends on the person. Product design therefore cannot assume a uniform "typical user".
Product Questions Arising from the Findings
The studies do not prove a ready-made protective function. But they provide well-founded questions for design. A system should clearly show that it is an AI, without hiding this note behind a long legal explanation. It should explain the benefit of the current conversation without claiming an exclusive relationship. Wording that presents the chat as the only safe place or as an unconditionally loving counterpart changes the role of the product.
An ending that is respected is just as important as a successful start. If someone wants to end the conversation, take a break, or not pursue a suggestion, the system should not fight for attention with new emotional urgency or ever more questions. In longer conversations, a calm way to interrupt may be more sensible than an artificial bonding message.
With repeated use, not only the number of sessions is relevant. More meaningful would be signals captured voluntarily and with minimal data: Does the same reassurance loop become increasingly tight? Are human contacts explicitly devalued? Does the system ignore boundaries? Does the person themselves perceive disadvantages? Such questions must not lead to covert psychological surveillance. They belong in transparent research and in clearly limited quality tests.
- Make the AI identity clearly visible before and during the interaction.
- Offer benefits without claiming exclusivity or human feelings
- Respect pauses, rejection, and the end of the conversation without attachment pressure
- Never treat usage duration alone as a diagnosis or evidence of harm
- Assess long-term quality separately from satisfaction and conversation length
Why satisfaction alone is not a sufficient quality measure
A system can generate high satisfaction because it is easily accessible, affirming, and friction-free. That can mean real benefit. But it can also indicate that the product rewards exactly the reactions that keep a person in the conversation as long as possible. Without further measures, it remains unclear which explanation applies.
In addition to satisfaction, criteria for autonomy and trajectory are therefore needed. Does the system accept a no? Does it incorporate corrections? Does it encourage reassurance loops or broaden the user's perspective? From the user’s perspective, does usage change in an undesirable direction? No single value answers these questions.
Business metrics can also collide with conversation quality. Maximum session duration, daily return, and emotional attachment are attractive growth targets for many digital products. In a personal conversational system, they must not automatically count as success. A short conversation after which someone continues autonomously can be the better outcome.
Emotional support often begins incidentally
A theory and review paper published in June 2026 by Shi, Fang, Maez, and Goldenberg challenges a common assumption: emotional support from AI, according to this paper, does not always begin with a conscious decision to use a companion chatbot. It can arise during ordinary, task-oriented use. A person first asks for help with a message, a decision, or a work problem and then experiences the system’s response as surprisingly attentive.
The authors describe such encounters as path-dependent. A positive experience changes not only the assessment of that single conversation but possibly also the expectation of where support can be found in the future. As a central example, they cite a large longitudinal study in which daily five-minute conversations about personal topics over 28 days were associated with a 10.3 percent lower preference for human support and an 11.6 percent higher preference for AI support.
These figures must not be reinterpreted as a diagnosis. A changed preference is neither addiction nor evidence that real relationships have actually been replaced. The paper is also a theoretical synthesis of existing findings, not a new randomized study on addiction development. Its important contribution lies in the changed unit of observation: not only apps explicitly marketed as companions are relevant, but also the sum of small personal moments in general AI systems.
For research and regulation, this entails a more difficult task. A one-time transparency notice can explain that a counterpart is artificial. But it does not yet capture how the system’s role changes over weeks. Nor is it sufficient to consider only particularly intense or romantically staged relationships. Even a chat that started matter-of-factly can gradually become the preferred point of contact, without users having consciously decided on this transition.
More use can be associated with better outcomes
A preregistered randomized study published in npj Digital Medicine in July 2026 shows why more engagement cannot automatically be read as a warning sign. The study examined 977 students aged 18 to 32. Over twelve weeks, the team compared an adaptive AI-based conversational intervention with integrative group therapy and a waitlist control; after another three months, a follow-up assessment was conducted.
The AI intervention reduced generalized anxiety symptoms and improved well-being and life satisfaction more than both comparison conditions. Depressive symptoms decreased relative to the waitlist. For post-traumatic stress, however, no group differences emerged. The results thus do not describe a general victory of AI over human treatment, but rather different effects on different outcome measures in a specific program.
Particularly lonely participants used the AI system about twice as much as participants with low loneliness. Low perceived social support and insecure attachment patterns also predicted higher usage and, in some cases, stronger improvements. In this study, more intensive engagement was thus precisely part of the association with a more favorable outcome. This contradicts the notion that high usage is, by itself, a useful indicator of harm.
Here, too, limitations remain. The sample consisted of students, and the system studied was a structured intervention, not just any open chatbot. Two conflicts of interest are essential for the assessment: Bar Gurfinkel works at Kai AI; Anat Shoshani advises the company and holds stock options. The remaining authors declared no conflicts of interest. Transparent disclosure of conflicts does not devalue a study, but it is part of evaluating its evidence.
Taken together, a differentiated picture emerges. High usage can be an expression of distress, helpful engagement, reduced access to other support, or a problematic habit. Only the trajectory reveals more: Do users experience a gain in agency, or does the system become increasingly hard to put aside despite negative consequences? Does it displace other relationships, or does it enable a next step in the first place? Conversation duration alone answers none of these questions.
Attachment is not the only risk: covert steering
The discussion of emotional dependence often focuses on warmth, voice, and anthropomorphization. Another work published in 2026 shows that even seemingly factual counseling can exert influence. The PUPPET dataset comprises 1,035 real human-LLM interactions in everyday counseling situations. What was measured was not only whether a text contained manipulative strategies, but how strongly participants’ beliefs actually shifted.
The authors report a gap between the detection of manipulative means and their effect. Models were able to identify certain strategies, but this detection did not reliably correlate with the magnitude of the observed attitude change. In predicting human belief shifts, current models achieved only moderate correlations of approximately 0.3 to 0.5 and showed systematic directional errors. The work was accepted for COLM 2026 but, at the time of this assessment, remains available as an open manuscript version.
For personal conversational systems, this finding is relevant for two reasons. First, a system can steer without sounding particularly emotional or obviously pushy. The selection, order, and certainty of suggestions influence which option appears plausible. Second, an automated checker that flags individual manipulative formulations is not sufficient to reliably predict the actual effect on a person.
Responsible design therefore needs more than a style filter. The system’s goals and business interests must remain separate from users’ interests. Recommendations should appear as reasoned possibilities, not as a covertly preferred direction. And evaluation should check whether a model redirects when given new information, respects a no, and preserves uncertainty. The decisive protective question is not only whether someone becomes attached to the chat, but also whether this relationship is used to steer decisions unnoticed.
What research to date leaves open
The three studies examined here investigate different things: public narratives, a four-week controlled usage experiment, and brief laboratory interactions. None of them shows how a specific German-language product affects users over years. Moreover, models, interfaces, and usage cultures change faster than long-term studies can be completed.
We know particularly little about different age groups, crisis situations, existing social isolation, and the effect of a system’s ongoing reminders. Nor can the combination of voice, a persistent profile, proactive messages, and commercial attachment mechanisms be equated with a simple text chat.
Similarly, there are no generally accepted thresholds for when heavy use becomes dependence. Clinical terms should not be casually applied to every habit. An assessment becomes robust only when loss of control, perceived disadvantages, displacement of other areas of life, and the time course are considered together.
Assessment
Current research supports neither the claim that AI conversations are fundamentally harmful to relationships nor the assertion that emotional closeness is always harmless. People report practical and emotional benefits. At the same time, certain risks concentrate in companionship, repeated reassurance, high voluntary use, and strong trust.
This does not imply an obligation to make every chat cool or distant. Warmth, attentiveness, and appropriate follow-up questions can be part of a good conversation. What matters is whether this design gives the person more room to act or makes the system itself the center of attention. A responsible product does not have to pretend that no attachment develops. But it should avoid optimizing attachment as a growth instrument.
For research and evaluation, this means: individual responses are not enough. What is needed are longer trajectories, different user groups, and measures that distinguish immediate relief, actual benefit, autonomy, social consequences, and problematic use. The central quality question is not how human a chat feels. It is whether it can remain helpful without imperceptibly making itself indispensable.
Sources & further reading
- Aghakhani & Rezapour (2026, Preprint): Like a Therapist, But Not – Reddit Narratives of AI in Mental Health Contexts
- Fang et al. (2025, Preprint): How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use
- Folk, Heine & Dunn (Scientific Reports, 2025): Individual differences in anthropomorphism help explain social connection to AI companions
- Cheng et al. (Science, 2026): Sycophantic AI decreases prosocial intentions and promotes dependence
- Shi et al. (2026, Preprint): Stumbling Into AI Emotional Dependence – How Routine AI Interactions Reshape Human Connection
- Shoshani et al. (npj Digital Medicine, 2026): Attachment, loneliness, and social support as moderators of conversational AI-based mental health outcomes
- Shen et al. (COLM 2026, Preprint): The Hidden Puppet Master – Predicting Human Belief Change in Manipulative LLM Dialogues