Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
Reviews spanning two years are grouped into themes
Sullivan and colleagues analyzed user-generated reviews of an AI companion app over a period of two years. For this, they used Latent Dirichlet Allocation, a method for automatically detecting recurring themes in larger volumes of text.
The method identified five positive and four negative themes. It does not provide a direct measurement of clinical efficacy and has only limited insight into the individual context behind a review. Its strength lies in revealing patterns across many freely formulated experiences.
App reviews are also selective. Particularly enthusiastic or upset users are more likely to write them. The results therefore show which interpretations occur and how they can be related, not how large their share is among all users.
The topic model also does not interpret meaning like a qualitative interview analysis. It groups linguistic patterns, which are then professionally labeled and assessed. The identified themes are thus a map of the material, not an automatic truth about the motives or effects of individual users.
Human-likeness is a double-edged signal
Perceived human-likeness was among the positive themes. A system that naturally picks up on language, remembers, and responds appropriately can more easily be experienced as a social counterpart. This may increase accessibility and emotional resonance.
Yet the more human a system appears, the more human standards are applied. If it forgets important information, responds inappropriately, or repeats standard phrases, the error does not seem merely technical. It can appear as a lack of attention or diligence.
Human-likeness thus does not simply raise quality. It raises expectations. Product design must decide which social signals the system can reliably carry, rather than simulating as many as possible.
Friendship and implausibility are closely linked
Another positive theme was the perceived friendship with the AI. Regular availability, personal address, and a continuous course of conversation can encourage friendship-like use.
At the same time, implausibility emerged as a negative theme. A chatbot can express closeness and in the next moment produce false facts, invented memories, or contradictory statements. The language of the relationship makes such breaks more visible.
A reliable companion therefore must not only sound friendly. It needs stable boundaries, correctable behavior, and a clear separation between memory and invention. Otherwise, the staging of friendship becomes an amplifier of technical weaknesses.
Emotional support can make privacy concerns more acute
Those who receive emotional support often share personal information. This is precisely why, alongside positive mental experiences, perceived privacy violations emerged as a negative theme. Closeness increases the value and sensitivity of the stored data.
A general privacy notice at the edge of the website is not sufficient for this situation. People need to understand whether conversations are linked to an account, used for model improvement, transferred to service providers, or stored permanently.
The more personal the companion appears, the easier control should be. Deleting, exporting, limiting reminders, and using without an account are not merely technical options. They determine whether emotional openness becomes a fair exchange.
The uncanny arises from almost successful closeness
The topic of creepiness describes a discomfort that is not identical to mere technical coldness. Often, something feels uncanny precisely when it appears very human and yet fails to respond humanly at a crucial point.
A bot that brings up intimate memories unasked, reacts with exaggerated emotion, or hints at possessiveness crosses a social boundary. Technically, the function may be intentional; it may be experienced as an intrusion.
Good design therefore does not require maximum closeness, but adjustable closeness. Users should be able to decide how personal the chat speaks, which memories it uses, and whether it independently brings up previous topics.
The default selection should remain restrained. Closeness can be deliberately expanded later; once a feeling of intrusion has been created, it is harder to repair. Especially with new users, the system should first learn which kind of address is pleasant, rather than staging personalization as a surprise. Settings must make these decisions visible and allow them to be fully reversed at any time without hidden consequences.
Everyday experiences demonstrate practical benefits
Ta and colleagues analyzed 1,854 Replika reviews as well as open-ended responses from 66 users. They found companionship, a non-judgmental space, encouraging messages, and helpful information when other sources were unavailable.
These findings explain the positive side of the ambivalence. People experience not just a technical simulation but concrete support in everyday situations. A conversation can temporarily reduce loneliness or give the courage to say something out loud.
The benefit does not have to be disputed in order to take limits seriously. On the contrary: precisely because the support can be relevant, data protection, credibility, and respectful distance deserve the same attention as tone and personalization.
Lonely users do not automatically have the same experience
The Danish study by Herbener and Damholdt shows that socially supportive chatbot use was associated with higher loneliness and lower perceived support in a small group of school students.
This group may weigh social signals differently than purpose-oriented users. An always-available companion may gain greater significance when other support is lacking. At the same time, outages, changes, or inappropriate closeness can have stronger consequences.
Product reviews should therefore not only consider average values. Particularly relevant are the experiences of people for whom the chat has become not merely entertainment, but an important part of daily support.
Good humanity remains limited
The thematic analysis shows no simple struggle between positive and negative attributes. Perceived humanity can enable emotional support and friendship. The same design can raise expectations, heighten privacy concerns, and seem uncanny when errors occur.
A responsible companion should therefore use human conversational qualities without claiming a human existence. It can speak attentively, casually, and personally. It should not feign feelings, create exclusivity, or retain control over data and closeness.
The best version of a chatbot is not the one that most successfully passes as a human. It is the one whose natural language remains useful while its origin, limits, and responsibility are clear enough at all times to become visible not only at the next breakdown.
Sources & further reading
- Yulia Sullivan, Serge Nyawa, Samuel Fosso Wamba (2023): Combating Loneliness with Artificial Intelligence: An AI-Based Emotional Support Model
- Vivian P. Ta, Caroline Griffith, Carolynn Boatfield (2020): User Experiences of Social Support From Companion Chatbots in Everyday Contexts: Thematic Analysis
- Arthur Bran Herbener, Malene Flensborg Damholdt (2024): Are lonely youngsters turning to chatbots for companionship? The relationship between chatbot usage and social connectedness in Danish high-school students