Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
A large Japanese survey shows the other side
Nakagomi and colleagues analyzed data from 14,721 Japanese adults from nationwide internet panel surveys conducted in December 2024 and January 2025. They distinguished the use of AI companions from other AI use.
Well-being was measured in three areas: life satisfaction, happiness, and sense and meaning in life. The use of AI companions was associated with higher values in all three areas. Other AI use showed weaker or inconsistent associations.
The large sample makes the result statistically interesting. Here, too, however, these are cross-sectional associations. Those who use a companion may differ from other people in many characteristics.
The association was not the same for everyone
The Japanese study examined social networks, support, and loneliness as moderating factors. For friendship-based social integration, a U-shaped pattern emerged: the positive associations were strongest at moderate levels of social connectedness and weaker at very high or very low levels of connectedness.
Particularly strong positive associations appeared at high levels of loneliness. This aligns with the assumption that people with unmet social and emotional needs might benefit more. However, it remains unclear which mechanism is responsible for this effect.
The attenuated associations at very low levels of integration are equally important. An AI companion may be most useful when sufficient social resources remain for a conversation to build upon. This interpretation must be tested in future research.
Selection and effect occur simultaneously
Those who are distressed or lonely have a stronger reason to use a support offering. As a result, more distressed individuals are found among users. If their condition later improves compared to what it would have been without use, the offering may still have been helpful. A single measurement cannot disentangle these processes.
Assessing efficacy requires at least longitudinal data, appropriate comparison groups, or experimental designs. It is also important to know baseline levels and usage intensity. Otherwise, a group with higher needs may be mistakenly taken as evidence of harm.
Conversely, a positive association must not be sold as a promise of efficacy. Higher well-being may be related to income, interest in technology, existing support, or other unobserved factors.
In addition, there is the diversity of products. A freely configurable role-playing chat, a structured mental health companion, and a general-purpose language model may appear under the same category, even though their behavior differs greatly. A claim about the efficacy of “AI companions” remains imprecise if the model, conversation design, memory, and purpose of use are not described.
App reviews reveal the same ambivalence
Sullivan and colleagues analyzed reviews of an AI companion app created over a two-year period using an LDA topic model. Positive themes concerned perceived humanity, emotional support, friendship, reduced loneliness, and mental benefits.
Negative themes included a lack of conscientiousness, insufficient credibility, privacy violations, and an uncanny feeling. The proposed model understands positive and negative features as interconnected.
This simultaneity aligns with the rest of the research. A system can provide relief in one moment and generate distrust in another. Average values obscure such shifting experiences.
Products need trajectory measures rather than moments of success
For a companion chat, it is not enough to ask after the conversation whether the response was helpful. A pleasant moment says little about whether loneliness, self-efficacy, or social integration change over weeks.
Meaningful observation could voluntarily capture whether users feel clearer, more connected, or more capable of acting after conversations. Equally important are signals that the chat is replacing human contact, reinforcing withdrawal, or has become an obligation.
Such data must be collected sparingly, transparently, and without manipulation. A system must not measure vulnerability merely to optimize engagement. Quality measurement should limit and improve the product, not classify people.
Another benchmark is behavior after a correction. If someone says they only want to talk or find a suggestion inappropriate, the chat must change course. Such small points in the conversation trajectory can be more telling of real conversation quality than a one-off general satisfaction rating.
The honest answer is: It depends on the course of the conversation
The Danish study shows that socially supportive chatbot use occurs among a small group of adolescents who are more lonely and less supported. The Japanese study finds positive associations between AI companions and several forms of well-being, particularly at high levels of loneliness.
Both findings are plausible, and both remain causally open. Neither “chatbots make people lonely” nor “AI companions increase well-being” is proven by these cross-sectional data.
For responsible development, this is not an unsatisfactory gap but a clear directive. What need to be examined are different starting conditions, actual usage patterns, and changes over time. Only then can it be assessed when a companion provides relief, when it is merely chosen, and when it may widen a social gap.
Sources & further reading
- Arthur Bran Herbener, Malene Flensborg Damholdt (2024): Are lonely youngsters turning to chatbots for companionship? The relationship between chatbot usage and social connectedness in Danish high-school students
- Atsushi Nakagomi, Yasuko Akutsu, Mika Yasuoka (2026): AI companions and subjective well-being: Moderation by social connectedness and loneliness
- Yulia Sullivan, Serge Nyawa, Samuel Fosso Wamba (2023): Combating Loneliness with Artificial Intelligence: An AI-Based Emotional Support Model