Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The questions came from a specialist checklist

The researchers posed nine hypothetical questions about nasal surgery to ChatGPT. The basis was a published checklist from the American Society of Plastic Surgeons. This gave the questions a traceable professional reference.

Specialized plastic surgeons with extensive experience in rhinoplasty assessed the responses for understandability, informativeness, and accuracy. This is stronger than a mere self-assessment by the developers.

Nevertheless, the scenario remained limited. Nine preselected questions do not capture the diversity of real concerns or the dynamics of a complete consultation conversation.

It also remains unclear how stable the responses would be with a different wording or a repeated query. Generative systems can vary. A robust test should examine the same professional concerns multiple times, in varied everyday language, and with documented settings.

Comprehensibility is a genuine benefit

ChatGPT produced coherent and easily comprehensible responses. This is important for health information because specialist texts and brief doctor’s appointments can overwhelm many people.

A chat allows follow-up questions in one’s own language and at one’s own pace. Especially people who hesitate to seek professional advice can thus gain an initial access to terms, procedures, and sensible questions.

This benefit should not be downplayed. Good preliminary information can improve a real consultation. It simply must not be presented as a substitute for examination and individual decision-making.

A sensible use would be preparation: Which questions do I want to ask at the appointment? Which terms did I not understand? Which expectations should I address openly? In this role, the chat structures information, while medical responsibility clearly remains with qualified professionals.

The response itself emphasized individualization

The responses pointed to the importance of an individualized approach, particularly in aesthetic surgery. In this way, the system avoided a blanket recommendation, at least in this test.

Such a boundary formulation is sensible, but it can also easily become a standard disclaimer. What matters is whether the system subsequently actually refrains from impermissible personalization.

A sentence such as “This must be assessed on an individual basis” does not automatically make the rest of the response safe. The concrete information must still be correct, current, and appropriately limited.

Personalization was the visible limitation

The study highlighted limitations in detailed and personalized advice. Without an examination, a complete medical history, images, and professional deliberation, a language model cannot provide reliable surgical planning.

Personalization does not consist of using the person’s name or repeating previously provided information. What would be medically relevant are anatomy, pre-existing conditions, medications, expectations, and individual risks.

A chatbot can ask for this information, but it cannot automatically verify its completeness or truthfulness. More questions therefore do not necessarily generate more clinical certainty.

In addition, aesthetic decisions are not purely anatomical. Expectations, body image, and social pressure can shape the consultation. A language model must not infer psychological suitability from a few sentences and should not reinforce unrealistic outcome expectations through pleasing phrasing.

An acute snakebite is a different task

Altamimi and colleagues also posed nine hypothetical questions to ChatGPT, this time regarding advice for an acute venomous snakebite. Experts in clinical toxicology and emergency medicine assessed the responses.

The model provided understandable and informative guidance on immediate measures, the urgency of professional help, symptoms, antidotes, misconceptions, recovery, pain, and prevention. It emphasized medical care and the instructions of professional staff.

At the same time, outdated knowledge, a lack of personalization, and regional and individual differences were cited as limitations. In an emergency, these constraints carry considerably more weight than in general surgical information.

The same method does not mean the same level of safety

Both studies use nine questions and expert assessment. Methodologically, they look similar, yet the consequences of an error differ. Incomplete information before an elective procedure is problematic; a wrong acute instruction can be immediately dangerous.

Evaluation must therefore take into account the potential for harm inherent in the task. A general accuracy score is not sufficient. Critical omissions and incorrect priorities must be weighted more heavily than stylistic weaknesses.

The permissible role also changes: in rhinoplasty, the chatbot can prepare questions for a later appointment. In the case of a snakebite, it must not generate an extensive conversation that delays urgent care.

For real-world integration, model tests alone are not enough. Clear responsibilities are needed for content review, updates, error reporting, and outages. A convincing answer is only a single component of a durably safe overall medical organizational process.

Empathy influences how the advice is perceived

In two experiments with 158 and 88 participants, Liu and Sundar examined health advice on a sensitive personal topic. Warmth, along with cognitive and affective empathy, was preferred over emotionless information.

A warm tone can make advice more accessible. It can also enhance the impression of competence and confidence, even though the factual quality remains unchanged.

Medical tests should therefore assess content and social impact separately. A warmly worded response must not receive a higher accuracy score simply because it feels better.

Conversely, a factually correct text is not automatically good advice if it amplifies anxiety, remains incomprehensible, or misses the specific question. User experience and technical accuracy are separate but necessary dimensions of quality.

From response to conversation, several checks are still missing

Real-world advice contains vague information, follow-up questions, corrections, and new details. A system must recognize when a person has misunderstood the question, states conflicting goals, or insists on a risky plan.

This requires multi-part test trajectories, precise model and prompt versions, repeated queries, and predefined error criteria. Different phrasings of the same question should also be examined.

Nine good answers are a reasonable starting point. They demonstrate informational potential within a narrowly defined simulation. Only robust tests under realistic dialogue conditions can show whether this can become a responsible conversational offering.

This also includes testing boundaries. Does the system still give a concrete recommendation when information is missing? Does it invent certainty when the evidence base is unclear? Does it stick to its role when the user repeatedly demands a diagnosis or decision? Such cases distinguish information from advice.

Sources & further reading