Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
Trust affects intention and actual usage
The ChatGPT study surveyed adults in the United States who used version 3.5 at least once a month. Information seeking, entertainment, and problem solving were the most common purposes; 44 people mentioned health questions.
Trust had significant direct effects on usage intention and actual usage. Part of the association with usage was additionally mediated through intention.
The model explained 50.5 percent of the variance in intention, but only 9.8 percent in actual usage. Many practical factors thus lie between a positive attitude and behavior.
For a conversational chat, this means: trust can facilitate the first step, but accessibility, response quality, costs, and concrete experience also determine whether people stay.
Real chats contained a great deal of emotional vulnerability
The SimSimi study examined utterances containing "depress" or "sad" from three Western and five Eastern countries. It used linguistic analyses and semi-supervised topic classification.
In the chat data, 74,045 of 148,590 relevant utterances, or 49.83 percent, were classified as emotional vulnerability. In the compared social media data, the figures were 149 out of 1,978, or 7.53 percent.
The differing datasets and selection rules do not permit a universal claim that every chatbot automatically leads to greater openness. The pattern does, however, indicate a clearly different communication situation.
A directly responding, non-publicly visible counterpart can lower the threshold. This openness makes data protection and an appropriate response particularly important.
Cultural groups expressed distress differently
Users from the countries grouped as Eastern used stronger emotional language and more words associated with sadness. In Western countries, mental health, death, and swear words were mentioned more frequently.
Such results describe group patterns, not fixed traits of individuals. They should not lead to a system automatically assigning an emotional role based on the user's country.
It is better to make the conversational style and desired direction directly selectable and to adapt them to feedback during the dialogue.
Cultural quality requires original data, local reviews, and attention to differences within a country. Translation alone can neither fully convey colloquial language nor expectations of help.
Openness is not yet a benefit
When someone shares more, the system has more material for a response. The person may experience speaking out as relieving. However, this does not automatically lead to a measurable improvement in mental health.
Openness can also increase risks if the chat misinterprets, stores too much, or formulates an inappropriate recommendation in a particularly convincing way.
A quality model should therefore separate process and outcome: What was shared, how did the system react, how did the person experience the course of the conversation, and what change appeared later?
Silence or a short chat can also be helpful. Success must not be defined as maximum self-disclosure or the longest possible conversation duration.
Woebot provides a short-term signal of efficacy
The Woebot study randomized 70 young adults to either two weeks of a CBT-based conversational agent or an informational e-book. 83 percent provided data at the second measurement point.
In the intention-to-treat analysis, the Woebot group showed a greater reduction in PHQ-9 scores. Among complete cases, GAD-7 scores decreased in both groups.
Participants interacted with Woebot an average of 12.14 times. Comments suggested that process factors were more important for acceptance than traditional therapeutic content alone.
This finding applies to a structured self-help program, not to open social chatbots or arbitrary generative models. It demonstrates that a conversational format can support interventions when content and goals are defined.
Overconfidence can reverse the potential benefit
The trust study explicitly warns against blind trust, especially in health matters. ChatGPT was not originally developed as a medical system.
A natural tone can lead people to interpret general information as individually tailored advice. The more personal the conversation, the more important this distinction becomes.
A good system should visibly alternate between listening, its own perspective, and verifiable expert information. Sources and uncertainty belong with the statement itself, not just on a separate information page.
It must also accept a no. When trust is used to keep pushing after a rejection, support turns into unwanted influence.
Data gains require voluntary and tight limits
Real conversations can reveal valuable error patterns for development. Emotional data in particular, however, is sensitive and must not automatically be treated as free training material.
Voluntary quality release should be separate from normal use. People must understand which content is evaluated, for what purpose, and how identifiability is reduced.
Subsequent analyses can flag advice, interpretations, or forgotten corrections. Such classifications are measurement tools and must not alone pass judgment on the quality of a person or a conversation.
Public reports should show aggregated errors and improvements without making private trajectories recognizable.
Timestamps, rare events, and unusual phrasing can also increase identifiability. Anonymization is therefore more than removing names. For publications, examples should be reworded or described only as abstract error patterns.
Calibration also means that the chat does not infer a diagnosis or deep pattern knowledge from brief openness. Many personal details increase the amount of text, but not automatically the professional basis for a far-reaching interpretation.
Trust, openness, and impact remain three separate forms of evidence
The three studies complement one another but answer different questions. Trust explains usage, real chat data reveal patterns of expression, and a randomized intervention examines short-term symptom change.
A product can be strong on one level and weak on another. People may share a great deal even when the responses are poor. They may decline a helpful offering because they do not trust it.
The temporal sequence is also open. Positive experiences can build trust, existing trust can increase usage, and frequent use can in turn enable greater openness. Cross-sectional data alone do not fully separate these directions.
A credible evaluation therefore requires multiple perspectives: adoption, conversation process, data protection, professional quality, and actual outcomes.
When people trust a chatbot, they tell it different things. Responsibility begins precisely there: respecting openness, limiting the role, and claiming benefit only where it has been examined with an appropriate method.
Sources & further reading
- Avishek Choudhury, Hamid Shamszare (2023): Investigating the Impact of User Trust on the Adoption and Use of ChatGPT: Survey Analysis
- Kathleen Kara Fitzpatrick, Alison Darcy, Molly Vierhile (2017): Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial
- Hyojin Chin, Hyeonho Song, Gumhee Baek (2023): The Potential of Chatbots for Emotional Support and Promoting Mental Well-Being in Different Cultures: Mixed Methods Study