Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

From 7,752 hits to 137 studies

The researchers searched MEDLINE, Embase, and Web of Science from the inception of the databases to October 27, 2023. A specialist librarian supported the search. Two people screened titles, abstracts, and full texts; 137 primary studies were included.

The studies examined evaluated the clinical accuracy of generative chatbots in health advice. 55 studies concerned surgical topics, 51 internal medicine, and 13 primary care. Frequently, the focus was on treatment, diagnosis, or prevention.

The systematic selection makes the review more robust than a loose collection of examples. However, its subject was the reporting quality of the studies, not a pooled efficacy of all chatbots.

The exact model version almost always remained unclear

136 of the 137 studies examined inaccessible, closed language models and did not report enough information to clearly identify the specific version. This means a central prerequisite for reproducibility is missing.

A model name alone is not sufficient. Providers change systems, safety rules, and default behavior. A response from March can differ from a response in October, even if the article states the same brand name.

Only 54 studies, or 39.4 percent, even mentioned the date of the query. Without the version and time, a result can hardly be replicated or meaningfully transferred to current operations.

Screenshots or complete raw responses would also be helpful. They allow other researchers to check whether the summary assessment matches the actual text. Where neither the response nor the exact system version is accessible, only the original study team’s judgment remains, with no robust independent verification.

Technical characteristics were inadequately described

All included studies insufficiently described model features such as temperature, token length, possible fine-tuning, layers, and other technical details. Some information is not even accessible to researchers for closed models.

This lack of transparency is not merely a formal issue. Temperature and output length can alter responses. Context windows, system instructions, and model updates influence whether a chat responds precisely, cautiously, or in detail.

When essential parameters remain unknown, a study examines less a clearly defined system than a time-limited access to a product. Its results can still be interesting, but they must be formulated with correspondingly narrow claims.

Prompt development was missing in 99.3 percent

136 studies described no phase of prompt engineering. This is remarkable because the formulation of a question to language models is part of the experimental setup. Even small differences can change the content, length, and tone of the response.

A fair evaluation does not need to optimize for a specific result over weeks. But it should explain how prompts were developed, whether multiple variants were tested, and whether the same rules applied to all systems under evaluation.

For conversational chats, a single question is not sufficient anyway. The system prompt, the previous conversation trajectory, user settings, and rephrasing after an error are all part of the product. Anyone who tests only an isolated prompt is not evaluating the complete conversational system.

There is also the risk of optimizing for known questions. If developers repeatedly use the same examples, the prompt may perform excellently on exactly those cases and fail on new conversation trajectories. Development cases and untouched test cases must therefore remain separate, and their origin must be documented.

Subjective success measures dominated

89 studies, or 65 percent, used subjective methods to define successful chatbot performance. Expert judgments are important in medical matters, but their criteria and inter-rater agreement must be transparent.

Terms such as helpful, complete, or empathetic can be understood in different ways. A response can be medically correct yet incomprehensible to laypeople. It can be worded accessibly and still omit a crucial safety note.

Good evaluation therefore separates dimensions: factual accuracy, completeness, comprehensibility, uncertainty, the limits of personalization, and potential risks. An overall judgment obscures why an answer passed.

Multiple independent evaluators and documented agreement can improve subjective measures. The assessment becomes even stronger when objective error criteria are added: fabricated sources, incorrect dosages, omitted warning signs, or a recommendation that should not be given without an examination.

Nine questions are not care provision

The supplementary rhinoplasty study posed nine questions from a checklist of the American Society of Plastic Surgeons to ChatGPT. Experienced specialists assessed accessibility, information content, and accuracy.

The answers were coherent and understandable and emphasized an individualized approach. At the same time, more detailed or personalized advice was lacking. The work classified itself as an observational study at evidence level V.

The result demonstrates potential information quality in a limited simulation. It neither proves safe advice in arbitrary cases nor behavior over longer dialogues. A good initial conversation on nine prepared questions is not clinical integration.

Empathy additionally alters the evaluation

Liu and Sundar examined medical chatbot consultations on a sensitive personal topic in two experiments with 158 and 88 participants. Sympathy as well as cognitive and affective empathy were rated more positively than purely emotionless advice.

The effect was particularly pronounced among people who were initially skeptical about whether machines possess social cognitive abilities. This makes it clear that tone influences the perception of health advice.

A methodologically sound study must therefore separate content from social presentation. A sympathetic response may be better received without being medically more correct. Conversely, a correct response can lose usability due to an inappropriate tone.

Reporting standards are part of safety

The review is intended to support the development of the Chatbot Assessment Reporting Tool, or CHART for short. Such a standard can specify which information studies should disclose regarding model, version, query date, prompts, evaluation, and safety.

The same principle applies to product developers. Model name, prompt version, test cases, and changes must be documented. Only then can it be traced why a result has improved or worsened and whether an update has introduced unexamined side effects.

137 studies sound like a broad evidence base. Without a reproducible experimental setup, a large portion of them remains hard to compare. Good research therefore begins not with the most convincing answer, but with a description that allows others to truly replicate the test.

Sources & further reading