Source-based15 min
Methodology
Conversation quality does not emerge in a single sentence. An overview of 125 studies and three benchmarks for multiple complete multi-turn conversation trajectories shows why technique, fit, relationship impression, responsibility, and trajectory must be examined separately.
Source-based11 min
Research review
A CHI 2026 study compares GPT-4o, GPT-5, and highly rated Reddit advice. The results are strong – but narrower than the headline suggests.
Source-based10 min
Analysis
Among 193 Italian millennials, a social conversational style increased perceived social presence, while the mere presence of an avatar did not. Trust arises more from behavior than from the surface face.
Source-based11 min
Analysis
An analysis of 152,783 utterances from eight countries found cultural differences and significantly more emotional vulnerability in chat conversations than in social media. The context shapes what people say.
Source-based11 min
Analysis
In 70 young adults, the PHQ-9 declined more strongly in the Woebot group than with an informational e-book. The brief study demonstrates feasibility and an initial signal—not the efficacy of arbitrary AI chats.
Source-based11 min
Analysis
Among 355 students, attitude, expected performance, social influences, and contextual conditions shaped usage. Anthropomorphism, novelty, and trust improved attitude—institutional rules changed the step to practice.
Source-based10 min
Analysis
For four Indian banking chatbots, trust explained measurable but limited portions of attitude, usage intention, and satisfaction. A trustworthy impression matters—but it is not a complete product model.
Source-based11 min
Analysis
A pilot with 17 evaluated adolescents showed high acceptance and a stronger PHQ-9 decline in the app group. The sample was too small for evidence of efficacy—this is exactly what the authors themselves state.
Source-based10 min
Analysis
In a study of 343 e-commerce customers, interactivity and perceived human-likeness influenced trust in chatbots. Trust mediated the intention to use them—an important product effect with a clear limitation.
Source-based11 min
Analysis
An overview categorizes LLM trustworthiness into seven main and 29 subcategories. Reliability, fairness, robustness, and social norms can develop differently—a single overall score is insufficient.
Source-based11 min
Analysis
A meta-analysis of 29 interventions found a small effect on psychological distress, but no significant effect on well-being. Platform, interaction type, and response generation influenced the outcome.
Source-based11 min
Analysis
GermanPartiesQA tests six commercial models with 418 statements from eleven German voting advice applications. The models had factual gaps and model-specific political patterns; persona adaptation was not automatically sycophancy.
Source-based11 min
Analysis
A large Brazilian study found small improvements in body image and well-being. At the same time, 61.9 percent in the intervention arm dropped out. Scalability and engagement must be assessed together.
Source-based11 min
Analysis
Trust influences usage, real chat data show unusually high emotional vulnerability, and structured offerings can help in the short term. These three levels must not be merged into a single promise of efficacy.