Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The study examines an offering for the ordinary postpartum routine

The lead publication from 2023 did not primarily examine a group with diagnosed depression. English-speaking women aged 18 and older who were to be discharged after a live birth from a tertiary academic center and who owned a smartphone were eligible to participate. They were enrolled during their birth hospitalization. The research question was correspondingly pragmatic: Is a digital conversational agent with perinatally oriented content acceptable in a general postpartum population, and do initial indications of an effect on mood and anxiety emerge?

This delimitation is crucial. Anyone who derives a statement about the treatment of pronounced postpartum depression from the study unknowingly shifts the target group. The average baseline values were below the thresholds for a positive depression screening. The study therefore primarily examined whether a low-threshold digital offering is accepted at all in a life phase that is physically, temporally, and organizationally exceptional, and whether measurable differences from usual care emerge within six weeks. This is not a minor preliminary stage, but an independent product and care question.

Randomization creates comparability, but not complete blinding

A total of 192 women were randomized in a one-to-one ratio. The intervention group installed a smartphone app with the conversational agent and perinatal content; in addition, they continued to receive usual care. The control group received only this usual care, consisting of routine follow-up care and the mental health care arranged by the respective obstetric provider. The study was unblinded. All participants completed questionnaires at baseline and at two, four, and six weeks after birth.

The primary endpoint was defined as the change in depressive symptoms after six weeks. For this, the research team used two established screening instruments, the Patient Health Questionnaire-9 and the Edinburgh Postnatal Depression Scale. Anxiety symptoms were captured with the Generalized Anxiety Disorder-7; satisfaction and acceptance were among the secondary outcomes. Of the 192 randomized women, 152 were included in the final analysis: 68 from the app group and 84 from usual care. This difference between allocation and analysis should remain visible in the interpretation, even if it does not render the randomized approach worthless.

One depression measure changes, the other does not

After six weeks, the average decline in the PHQ-9 was larger in the app group than in the control group: 1.32 points versus 0.13 points. This is a positive signal in favor of the intervention. However, it does not by itself represent the entire primary endpoint, because on the Edinburgh Postnatal Depression Scale, the group means did not differ after six weeks. Nor was there a group difference on the GAD-7; the anxiety values and the symptoms measured with the EPDS remained similar to baseline.

Thus, the finding is neither consistently negative nor consistently confirmatory. Two instruments for depressive symptoms lead to different results. The provided source references do not report a confidence interval or a p-value for the PHQ-9 comparison, so its statistical precision cannot be further assessed here. From an editorial standpoint, we consider it wrong to downplay the PHQ-9 decline simply because it occurs in isolation. It would be equally wrong to present it as comprehensive evidence of efficacy. The appropriate finding is more narrowly stated: In a low-burden general population, a greater decline was seen on one depression measure, but not on the other symptom measures.

Satisfaction and actual use tell different stories

The acceptance data are more impressive than the symptom picture. Of the users included in the final analysis, 91 percent, or 62 women, described themselves as satisfied or very satisfied. Eighty percent of all study participants stated that they felt comfortable with a mobile app for mood regulation. At the same time, 74 percent of the intervention group reported having used the conversational agent at least once within the two weeks before the final survey.

These figures must not be merged into a single scale of agreement. Feeling comfortable with the format, being satisfied with it, and currently using it are different matters. Using it at least once in two weeks does not demonstrate high intensity, but it does show that the offering did not immediately become meaningless for a considerable proportion of the women after the initial installation. Our assessment is therefore explicitly positive: In a phase in which time, availability, and costs can make access to support difficult, a broad willingness to use it is a substantial finding. It just does not automatically prove an effect on symptoms.

Low baseline values limit the possible decline

The authors themselves identify the central limitation: at the outset, the sample on average fell below the thresholds of a positive depression screening. Where baseline burden is low, there is less room for marked improvement on a symptom scale. This so-called floor problem is a plausible explanation for the limited differences, but it is not retrospective evidence that the app would have worked better in more severely affected women. That is precisely what another study would need to examine.

The finding also changes the meaning of the word 'effect.' In a general obstetric population, a digital conversational offering can have value without producing large average symptom shifts—for example, as an easily accessible format that women accept and repeatedly access. Whether this yields prevention, earlier help, or some other long-term benefit was not demonstrated in the six weeks. The study thus provides good reasons for subsequent experiments with a longer duration and targeted inclusion of risk groups, as the publication itself suggests. It provides no basis for simply transferring the observed benefit to these groups.

The Argentine pilot study reveals the same methodological conflict

A randomized pilot study by María Carolina Klos and colleagues examined the AI-based conversational agent Tess among Argentine university students. 181 people between 18 and 33 years of age were assigned either to Tess for eight weeks or to a psychoeducational book about depression; 87.2 percent of the initial sample were women. After eight weeks, however, data were available for only 39 of the 99 people in the intervention group and 34 of the 82 people in the control group. This low completion rate substantially limits the conclusions that can be drawn.

No significant differences were found between the groups for depressive symptoms or anxiety. Within the Tess group, anxiety symptoms decreased, but not in the control group. For causal assessment, however, the direct group comparison is decisive; a before-and-after difference within only one group does not replace it. At the same time, the students used the system intensively: on average, 472 messages were exchanged, of which 116 were responses from the users. More messages were associated with more positive feedback. Here too, the direction is open—more satisfied people might have written longer, or longer use might have improved the evaluation. Compared with the postpartum study, a serious pattern nevertheless repeats itself: usage and agreement are easier to demonstrate than clear symptom benefits.

The meta-analysis supports small effects, not boundless promises

The broader body of research is more favorable than two individual trials might suggest. Jake Linardon and colleagues evaluated a total of 176 randomized studies on smartphone apps for depressive and anxiety symptoms in 2024. Across 33,567 participants, apps showed a significant, overall small effect on depressive symptoms with a standardized effect size of g = 0.28. For generalized anxiety, the likewise small effect was g = 0.26 and was based on 22,394 participants. The effects remained stable at different follow-up time points and after excluding smaller studies and studies with a higher risk of bias.

For depression, effect sizes were larger when apps included elements of cognitive behavioral therapy or conversational technology. For anxiety, they were larger, among other things, when generalized anxiety was the primary target and the app offered corresponding behavioral therapeutic or mood-monitoring functions. This usefully extends the individual finding: digital applications can, on average, achieve measurable, albeit small, symptom improvements. But it does not allow the conclusion that the conversational form alone causes the additional effect. The features are associated with larger effects in the meta-analysis; they were not randomized as isolated components. Moreover, an average across many apps says little about whether a specific product is suitable for a specific postpartum group.

The right standard is a clearly limited task

The three publications do not yield a judgment about AI conversations in general. The postpartum study examined a specific smartphone application with perinatal content; the provided information does not describe its technical architecture in detail. Results about this conversational agent therefore cannot be readily transferred to open generative systems. Conversely, it would be equally premature to infer from an inconsistent primary outcome that digital conversations are clinically irrelevant. The meta-analysis shows that small average effects can be robust across a large number of randomized trials.

For product development and research, a more precise guiding question follows: What limited task should a conversational system fulfill in which population over what time period? In the study by Suharwardy and colleagues, the most convincing contribution was not a dramatic reduction in symptoms, but rather the combination of high satisfaction, continued use, and a single positive depression signal. This is less spectacular than a promise of comprehensive care, but more professionally useful. Our editorial position is therefore clear: Acceptance is not a surrogate endpoint for efficacy, but for low-threshold offerings it is also not a consolation prize. It is an independent outcome that can justify a meaningful next study – provided that the limited symptom evidence is disclosed just as openly.

Sources & further reading