Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
Article history
- First published
The short path from a good conversation to a big promise
Voice-based systems have an obvious appeal: they can offer information in a familiar interaction format, accept questions, and provide repeated support. This quickly creates the idea of a scalable health offering. Yet between an understandable answer, a pleasant user experience, and an improvement in health-relevant outcomes lie several levels of evidence. It is precisely the conversational format that tempts one to skip them mentally. Anyone who feels addressed may experience an application as helpful, even though neither a behavior change nor a health effect has been demonstrated.
This distinction is not a devaluation of perceived quality. Acceptance plays a role in determining whether a digital offering is used at all. Likewise, correct speech recognition is not a trivial matter, but a prerequisite for any further function. It becomes editorially problematic only when such findings stand in for efficacy. The key question is therefore not whether a voice is good or bad. What is decisive is which conclusion a study may draw from which measured value.
The systematic review primarily examines the type of evaluation
Bérubé and colleagues did not want to simply condense the then-current state of efficacy into a single effect number. Their systematic literature review examined the methods by which voice-based conversational agents for the prevention or management of chronic and mental illnesses were empirically evaluated. Primary research was included that examined at least system accuracy, technology acceptance, or both. Searches were conducted in PubMed MEDLINE, Embase, PsycINFO, Scopus, and Web of Science. Two people independently performed screening and data extraction; their agreement was recorded with Cohen's kappa. The selected works were narratively synthesized.
This design is central to the interpretation. The review maps a research field whose studies, by virtue of the inclusion criteria alone, did not necessarily have to examine health endpoints. Of 7170 publications screened in advance, twelve met the criteria. All included studies were non-experimental. Thus, the review was able to describe how systems functioned and were received, but could not obtain robust causal evidence on whether their use changed health or behavior. This is not a retrospectively discovered weakness of a single product, but the most important structural finding of this literature.
Twelve studies represent very different tasks
Under the common label of the language-based agent, different functions were grouped. Five systems offered behavioral support, three served health monitoring, and four combined both. The applications ran on smartphones, tablets, or smart speakers; in two cases, no device was specified. The diseases targeted were also broadly varied. Three systems dealt with cancer, and others addressed diabetes, heart failure, hearing impairment, asthma, Parkinson's disease, dementia, autism, intellectual disability, and depression, among others.
This breadth prevents a simple overall judgment. A system intended to answer health-related questions must be assessed differently than one that supports behavior or monitors data. Likewise, a finding from one disease field cannot be readily transferred to another. The twelve studies therefore initially demonstrate that language was tested for various health tasks. They do not demonstrate that a uniform intervention class with a common effect had already emerged from this. The heterogeneity here is not only a statistical challenge but a conceptual problem.
The measurements focused primarily on the technically straightforward.
Seven of the twelve studies examined technology acceptance, but only three used validated instruments for this. Six studies reported performance measures for speech recognition or the system's ability to answer health-related questions. In contrast, only two studies captured behavior or attitudes toward the health behavior that the intervention targeted. Only four studies controlled for participants' prior technology experience. The risk of bias differed considerably between the studies.
This is where the core conflict lies. Speech recognition and answer performance are relatively directly observable at the system level; health-relevant behavior, on the other hand, arises in a social and temporal context. Acceptance can be assessed after a usage situation, whereas a robust effect requires a suitable comparison and an observation appropriate to the goal. Moreover, because all included studies were non-experimental, positive evaluations cannot be clearly attributed to the system as the cause. Prior familiarity with technology could also influence perception and use, but was only considered in a minority of the studies.
The positive signals are nevertheless real findings
The review assesses the results on system accuracy and technology acceptance as encouraging. That should not be downplayed. A speech-based offering that does not reliably process utterances or is rejected by its target group fails before it can have any possible health effect. The studies thus show that such systems were not merely theoretical concepts: they were practically deployed and were able to deliver relevant technical and usage-related signals in the contexts examined.
Our editorial assessment is therefore neither euphoria nor debunking. Acceptance is an independent developmental finding, but not a surrogate endpoint for health. Technical accuracy is a necessary performance dimension, but it does not answer whether an intervention has additional benefit compared with usual care. The lead review also identifies exactly this comparison as an open task. Its judgment that the field is still in an early stage follows from the small, heterogeneous study base and the partly high risk of bias—not from evidence that speech-based support is fundamentally ineffective.
The psychiatric literature sounds more optimistic
The 2019 review by Vaidyam and colleagues broadens the perspective to conversational agents in psychiatric contexts. After a systematic search in June 2018, eight studies from the databases and two additional ones via reference lists were included. Applications for people with mental illnesses or increased risk were considered, including depression, anxiety, schizophrenia, bipolar disorders, and substance use disorders. The authors reported high potential, especially for psychoeducation and adherence. Satisfaction scores were also high in the included studies.
These findings support that good user experiences were not limited to the twelve studies of the lead review. At the same time, the language of potential remains crucial. High satisfaction can mean that people use an offering willingly and possibly regularly. By itself, it proves neither diagnostic quality nor treatment effect. Vaidyam and colleagues also described the evidence as preliminary and pointed to heterogeneous studies and the need for standardized outcome reporting. The more optimistic tone therefore contradicts the lead review less than it initially seems: both see constructive early signals, but no completed evidence base for efficacy.
Speech analysis answers a different question than conversation management
Le Glaz and colleagues examined machine learning and natural language processing in the field of mental health in 2020. Their systematic review, reported according to PRISMA and registered with PROSPERO, included 58 of 327 identified articles and qualitatively assessed the results. The studies primarily used medical records and social media. Their goals included extracting symptoms, assessing severity, comparing treatment outcomes, and gaining psychopathological clues. Powerful classifiers were preferred more often than models with transparent operation.
For conversations with AI, this literature is relevant but cannot be equated with intervention research. A model can classify language without conducting a dialogue with a person or offering support. The review also describes social media as an imprecisely defined cohort and points out that language-specific features can influence performance. Its authors see potential in making previously little-exploited data on everyday habits usable, but they view the procedures as support for clinical practice and identify open ethical questions. Analytical capability, conversation quality, and health impact thus remain three distinct objects of examination.
The claim must correspond to the task that was tested
The three reviews do not yield an overall judgment of today's AI systems; their search periods, subject areas, and methods are too limited for that. However, they do provide a reliable rule for evaluation. If a system is to recognize speech or answer health-related questions, accuracy and error profiles are appropriate metrics. If it is to be accepted and used repeatedly, it needs comprehensible acceptance measures. If it is to change behavior or health, exactly those outcomes must be collected and examined in suitable comparative designs. None of these levels may automatically vouch for the next.
Our editorial position is therefore deliberately narrow: voice-based health assistants deserve neither a blanket promise of care nor reflexive disparagement. The state of research at that time showed technical feasibility and frequently positive reception more clearly than health efficacy. That is a useful basis for targeted development, provided that claims stay at the level of the evidence. Scientifically, the most convincing voice is not the one that sounds particularly human. It is the one whose role is precisely described and whose success was measured against exactly that role.
Sources & further reading
- Caterina Bérubé, Theresa Schachner, Roman Keller (2021): Voice-Based Conversational Agents for the Prevention and Management of Chronic and Mental Health Conditions: Systematic Literature Review
- Aditya Vaidyam, Hannah Wisniewski, John Halamka (2019): Chatbots and Conversational Agents in Mental Health: A Review of the Psychiatric Landscape
- Aziliz Le Glaz, Yannis Haralambous, Deok-Hee Kim-Dufor et al. (2020): Machine Learning and Natural Language Processing in Mental Health: Systematic Review