Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
Article history
- First published
70 young adults were randomly assigned
The participants were between 18 and 28 years old and were recruited online via a social media page from the university setting. 34 were assigned to the Woebot group, 36 to the information control.
The experiment was unblinded: participants knew whether they were working with a chatbot or an e-book. Expectations of a new interactive product may therefore have influenced part of the experience.
Depression was measured with the PHQ-9, anxiety with the GAD-7, and positive and negative affect at baseline and approximately two to three weeks later.
83 percent provided data at the second measurement point. The 17 percent dropout rate is relevant for interpretation but was addressed through an intention-to-treat analysis for the central depression comparison.
The PHQ-9 changed between the groups
At baseline, there were no significant group differences. At the second measurement point, the intention-to-treat analysis showed a significant difference in depression: the Woebot group reduced their PHQ-9 scores, while the information group did not to a corresponding degree.
The result is a preliminary signal of efficacy for the combination of conversational format and CBT-based self-help program under investigation. It does not indicate which individual component triggered the change.
Possible factors include regular use, immediate feedback, the division into small units, and the impression of a responsive counterpart. A comparison with the same content in a non-dialogic app would have isolated the format effect more precisely.
Two weeks show short-term change. Whether the difference persisted, whether further use was necessary, or how the offering should be embedded in care cannot be derived from this.
Both groups improved on anxiety
In the analysis of participants with complete data, GAD-7 scores decreased significantly in both groups. For anxiety, no pattern comparable to that of depression emerged.
An informational e-book can also provide orientation and relief. Moreover, repeated measurements, time, or other events during the short period can influence scores.
The differentiated result precludes the simplified claim that the chatbot reduced all examined symptoms more effectively. Each scale and analysis population must be reported separately.
For product communication, this means not deriving a comprehensive promise about mental health from a single positive primary finding.
Use remained comparatively high for two weeks
Participants in the Woebot group interacted with the system an average of 12.14 times. Given the intended duration of two weeks, this indicates regular use.
Digital self-help often fails not only because of content but also because of low engagement. A conversational format can make small units more accessible and make it easier for users to return.
High activity in a short study, however, says little about months-long use. Novelty, study contact, and a clearly limited task can initially increase engagement.
Long-term evaluation should examine which individuals return, who drops out early, and whether frequent use is associated with benefit or possibly also with dependence on the application.
Process factors shaped acceptance
In the comments, process factors appeared to be more important for acceptance than content that reflected traditional therapy. This is a central lesson for conversational products.
Expertly sound content can go unused if the interaction feels rigid, repetitive, or instructive. Conversely, a pleasant process must not mask professional weaknesses.
Process quality encompasses timing, understandable language, freedom of choice, and responsiveness to feedback. Young adults in particular may want support without every message sounding like a therapy session.
Combining professional structure with natural dialogue is therefore a development task in its own right. It cannot be guaranteed by an extensive system prompt alone.
The transition between units is also part of the process. A system should not pretend that each session starts from scratch, but it should incorporate earlier content only with agreement and a comprehensible storage logic.
A youth pilot was feasible but too small for evidence of efficacy
Nicol and colleagues studied a CBT chatbot in 13- to 17-year-olds with moderate depressive symptoms in primary care. Of 18 randomized adolescents, 17 were included in the analysis.
After four weeks, the mean PHQ-9 score decreased by 3.3 points in the app group and by two points in the waitlist group. Adolescents and parents rated feasibility, usability, and acceptance highly.
The authors explicitly emphasize that this small pilot could not establish evidence of efficacy. Its purpose was to determine whether a larger study would be feasible.
The predominantly female, white, and privately insured sample limits generalizability. Subsequent research must include rural, socioeconomically disadvantaged, and underrepresented groups.
A large Brazilian study found small effects and high attrition
Matheson and colleagues randomized 1,715 Brazilian adolescents aged 13 to 18 to a body-image chatbot or a survey-only control condition. The study was fully remote and preregistered.
The chatbot offered brief micro-interventions over 72 hours. On several primary and secondary measures, small significant improvements emerged, including state and trait body image as well as individual affect and self-efficacy measures.
Of the 858 participants in the intervention arm, 531 were lost, an attrition rate of 61.9 percent. Among the 327 who entered the chatbot, 258 completed at least one technique and used five on average.
Scalability and small positive effects thus stand alongside a major engagement problem. Reach only translates into impact if people actually access the offering and use it sufficiently.
The right conclusion is specific rather than sweeping
The Woebot study provides evidence of a short-term positive signal for a specific CBT-based conversational program in a small group of young adults. It does not provide evidence that free-form generative chats have the same effect.
The supplementary adolescent studies show how strongly population, content, duration, and implementation alter the results. Feasibility, acceptance, symptom change, and actual use are distinct questions.
For current systems, larger, more diverse, and longer studies are needed, along with precise documentation of the model, prompt, content, and safety architecture. Conversation errors should also be recorded as a separate outcome dimension.
Two weeks of Woebot was an important early step. It remains credible if the result is neither downplayed nor inflated into general evidence of AI support.
Sources & further reading
- Kathleen Kara Fitzpatrick, Alison Darcy, Molly Vierhile (2017): Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial
- Ginger E. Nicol, Ruoyun Wang, Sharon Graham (2022): Chatbot-Delivered Cognitive Behavioral Therapy in Adolescents With Depression and Anxiety During the COVID-19 Pandemic: Feasibility and Acceptability Study
- Emily L. Matheson, Harriet Smith, Ana Carolina Soares Amaral (2023): Using Chatbot Technology to Improve Brazilian Adolescents’ Body Image and Mental Health at Scale: Randomized Controlled Trial