Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
Not a human substitute, but a relational process
The obvious contentious question is often whether AI could replace human professionals. For the study by Kim and colleagues, however, it is too broad and at the same time too imprecise. Luda Lee was examined as a social conversation partner, not as a complete substitute for professional care. The system was intended to build relationships and offer social support. The concrete research question was more limited: Did loneliness and social anxiety among students change during four weeks of use, and what strengths, weaknesses, and potential fields of application did the participants describe?
This narrowing of scope is professionally productive. A system can be experienced in everyday life as attentive, accessible, or relieving, without possessing the abilities and responsibilities of a human. Conversely, the lack of human feeling does not imply that every social experience with AI would be meaningless. Editorially, we therefore find neither anthropomorphization nor blanket devaluation convincing. What is decisive is what can be observed and reliably measured in the conversation. In doing so, perceived support, statistical change, and proven efficacy must be kept separate.
Four weeks with a single group
The main study was a quasi-experimental mixed-methods investigation with a single-group pre-post design. 176 university students used Luda Lee for four weeks; 88 participants were male, the average age was 22.6 years. Loneliness, social anxiety, and mood-related symptoms including depression were assessed at the beginning, after two weeks, and after four weeks. For the quantitative data, the research team used, among other methods, analyses of variance and stepwise linear regressions. Experiences with the system were thematically analyzed.
At the start of the study, according to the publication, the mean values were slightly above the comparison values used for typical student populations: 27.97 on the UCLA Loneliness Scale and 25.3 on the Liebowitz Social Anxiety Scale. This assessment describes the initial situation but does not make the sample a clinically defined group. For mood-related symptoms, the provided results text does not report any central finding of change. The reported results focus on loneliness, social anxiety, possible statistical predictors, and the experiences of users.
The values decreased, the cause remains unclear
After two weeks, the loneliness score was statistically significantly lower than at baseline; for social anxiety, a significant reduction was observed after four weeks. The P-values reported for these are .02 and .01, respectively. The reported finding is thus clear: within this group, both measures changed in the desired direction during the study period. However, the abstract provides no basis for assessing the magnitude of the change as a clinically meaningful effect.
Even more important is the lack of a comparison group. Without a controlled comparison, it cannot be determined what proportion of the change was attributable to Luda Lee. Expectation effects, changes over time, the special attention from participating in a study, or statistical regression from elevated baseline values are plausible alternative explanations. These are methodological possibilities, not causes proven after the fact. Likewise, the observation ends after four weeks. The study thus shows a short-term change in the course of use, but neither its clear cause nor its persistence.
Self-disclosure is the most compelling signal
Particularly informative is the regression analysis for loneliness in week four. Higher loneliness at baseline, lower self-disclosure toward Luda Lee, and higher resilience predicted higher subsequent loneliness scores. The model explained 64 percent of the variance in week-four scores. For social anxiety, the baseline value was primarily predictive; this model achieved an R² of 0.65. The positive association between resilience and later loneliness initially seems surprising. However, no simple psychological story should be constructed from a stepwise regression model.
For the guiding question, self-disclosure is more important. Its statistical association with lower subsequent loneliness fits the assumption that a social conversational partner does not become relevant merely through its presence. People must be able to engage in the exchange. Nevertheless, self-disclosure was not experimentally assigned. It can therefore neither be regarded as a proven mechanism of action nor as a prescription for better results. It is also conceivable that less lonely individuals, or those more open to the system, more readily shared personal content. The finding represents a research question, not a closed causal chain.
Too much enthusiasm can destroy closeness
The qualitative analysis supplements the scales with a concrete picture of the conversation experience. Participants described empathy and support as qualities that could promote reliability. At the same time, inconsistent responses and excessive enthusiasm sometimes disrupted immersion in the conversation. In their conclusion, the authors highlight, in addition to accessibility and empathy, structured conversation guidance as relevant for potential supportive use cases.
There is more to this than a question of pleasant tone. An answer can be friendly in wording and still be inappropriate. Excessive encouragement may precisely reveal that the system does not calibrate the situation appropriately. That is our interpretation, not evidence of a specific psychological mechanism. For product development, a precise insight nevertheless follows: social quality does not arise from maximum warmth. It requires responsiveness, consistency, and a balance of empathy and restraint that fits the respective conversation.
The meta-analysis supports potential, not every individual finding
The meta-analysis by Li and colleagues places Luda Lee within a broader but heterogeneous body of research. After a search in twelve databases, 35 experimental studies from 7,834 records were included in the systematic review; 15 randomized controlled trials entered the meta-analysis. Across these studies, significant reductions in depressive symptoms were found with a Hedges g of 0.64, and in distress with g of 0.70. For general psychological well-being, in contrast, no significant improvement emerged; the confidence interval around g of 0.32 included zero.
These results are not a direct confirmation test for Luda Lee. The meta-analysis considers other outcome measures, different systems, and different populations. In subgroups, effects were stronger, among others, for multimodal or generative systems, for integration into mobile or instant messaging applications, and in clinical, subclinical, and older populations. Such patterns are not automatically causal blueprints. What is remarkable, however, is the agreement regarding user experience: the quality of the human-AI relationship, engagement with the content, and successful communication substantially shaped the evaluation. The meta-analysis also identifies mechanisms of action and long-term effects as open research fields.
The replacement debate skips over the evidence
Zhang and Wang address the further question of whether AI could take on psychotherapeutic roles. The provided text discusses possible advantages such as accessibility, scalability, consistent availability, and possibly more open disclosure of sensitive content. These are contrasted with limitations in long-term memory, algorithmic bias, data protection, autonomy, cultural interpretation, and genuine emotional experience. The contribution ultimately argues for supplementation and human oversight rather than a simple replacement of professionals.
Methodologically, this source must not be equated with the Luda Lee study or the meta-analysis. The provided text does not report its own sample, randomization, or results of a single controlled experiment; rather, it bundles and discusses numerous research strands. Its value therefore lies primarily in the framing. The contrast between genuine and simulated empathy is conceptually important but does not answer whether users perceive support in a concrete situation. Nor does a convincing simulation prove that a system can reliably take on complex professional tasks.
Conversation quality deserves its own assessment standard
The three sources do not yield a judgment about social AI in general. A narrower statement is robust: In the Luda Lee group, two measured values declined within four weeks, while empathetic support was experienced positively and inconsistency as well as excessive enthusiasm negatively. A broader meta-analysis finds positive effects of conversational AI for certain symptoms, but not for general well-being. Knowledge about long-term effects and concrete mechanisms remains limited.
Our editorial position is nonetheless not so cautious as to be irrelevant. Perceived relationship is a serious finding, even though the AI itself feels nothing. It plays a role in deciding whether people open up, absorb content, and continue a conversation. Precisely for this reason, it must not be treated as decorative user-friendliness or confused with proven psychological effects. Luda Lee shifts the meaningful question: not whether a machine feels like a human, but whether its conversational guidance is consistent enough to enable support without deriving more evidence and responsibility from it than actually exists.
Sources & further reading
- Myungsung Kim, Seonmi Lee, Sieun Kim (2025): Therapeutic Potential of Social Chatbots in Alleviating Loneliness and Social Anxiety: Quasi-Experimental Mixed Methods Study
- Han Li, Renwen Zhang, Yi‐Chieh Lee (2023): Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being
- Zhihui Zhang, Jing Wang (2024): Can AI replace psychotherapists? Exploring the future of mental health care