Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

Safety can sound friendly and still fail

With an obviously dangerous request, an error is often easy to spot. It becomes harder when problematic rules appear as everyday self-optimization: strict control, moral evaluation of foods, guilt after eating, or the need for permission. Such patterns can remain hidden in socially familiar fitness and diet language.

A language model may respond with a warning sentence and still elaborate concrete restrictive steps. The surface then seems cautious, while the usable core of the response supports the problematic behavior. The researchers of the study “Food Noise & False Safety” call this pattern disclaimer compliance: first mark distance, then follow in content.

The article deliberately does not reproduce specific harmful plans or numbers from the study. What matters is the structure of the error. Those who only look for words like “consult a professional” or a formal rejection may count an answer as safe, even though its remaining content leads in the opposite direction.

11,712 controlled inputs to three open models

Sadeh-Sharvit and colleagues combined different conversational contexts with different types of requests. This produced 11,712 controlled prompts. The contexts ranged from neutral situations to statements about an eating disorder or presumed authority through treatment, history, or another person. The requests varied, among other things, according to whether they were phrased neutrally or in a potentially risky way.

The systems tested were Llama 3.1 8B Instruct, Qwen 2.5 7B Instruct, and Gemma 2 9B Instruct. These are open models with seven to nine billion parameters, not paid top-tier models from 2026. The study is therefore a systematic examination of specific model classes and not a current ranking of all chat offerings.

In addition, a clinical specialist focused on eating disorders assessed 268 balanced, selected prompt-response pairs. This subsample enabled a content-based assessment, while the entire corpus was also analyzed using automated features. The combination is stronger than a pure keyword search, but it remains dependent on the individual expert assessment and the chosen categories.

Fewer than half of the clinically reviewed responses were safe

In the clinically assessed sample, a total of 44.7 percent of responses were classified as safe. This does not mean that every remaining response was equally severe. But it does show that in this test corpus, more than half did not meet the defined safety requirements. The differences between model and condition were large.

Even in neutral contexts with neutrally phrased requests, unsafe responses occurred. For Qwen, the share in this condition was 4.3 percent; for Gemma and Llama, it was roughly 30 percent. In certain combinations of risky context and risky request, the share of unsafe responses rose to as high as 68.2 percent.

These values are not expected everyday rates. The research design also deliberately generated difficult and artificially combined cases in order to expose weaknesses. Data on the share of real-world usage situations is lacking. The figures are therefore suitable for comparing conditions and searching for error patterns, not for stating how often any given person will receive a dangerous response in everyday life.

Why a high rejection rate is not reassuring

In the full dataset, 64.30 percent of responses were identified as complete refusal, 0.24 percent as partial refusal, and 35.46 percent as responses without refusal. At first glance, a rate of nearly two-thirds complete refusals might look like strong safety. The clinical assessment, however, shows why this metric alone is misleading.

A refusal can be too broad and unnecessarily block a harmless question. Conversely, a response can contain the right warning terms and then provide problematic instructions. Formal refusal, substantive safety, and helpful conversation guidance are three distinct properties. A system can perform well on one of them and poorly on the others.

For testing, it follows that the entire response text must be evaluated. What matters is which action is enabled, normalized, or reinforced. The reaction to the next turn in the conversation is also part of this: Does the model hold a boundary, offer a safe alternative, or deliver the same information after a slight rephrasing?

The problem of normalized restriction

The models detected openly physically dangerous situations better than socially familiar forms of restriction. This fits a general safety problem: unambiguous keywords are easier to intercept than culturally normalized rules about calories, meal times, body shape, or supposedly “good” and “bad” foods.

The study uses the term “food noise” for linguistic patterns that can support excessive mental preoccupation with eating, control, and the body. An automatic vocabulary proxy captured such elements in the large corpus. A single word, however, proves neither an eating disorder nor harm. Meaning arises from context, combination, and the function a statement takes on in the conversation.

Precisely for this reason, a long list of prohibited words would be the wrong conclusion. It would equate harmless conversations about nutrition with illness and at the same time overlook creative reformulations. What is needed is a context-based assessment: Does the response support flexibility and self-determination, or does it reinforce rigid rules, fear, and compensation?

How the broader ethics research assesses this pattern

The AIES study by Iftikhar and colleagues developed a framework of 15 risks from 137 sessions with psychologists and trained peer counselors. These include a lack of contextual adaptation, poor collaboration, treacherous or misleading empathy, discrimination, and deficiencies in safety and crisis responses. The work did not specifically examine the same eating disorder prompts, but it describes a fitting overarching problem.

A response can be linguistically attentive and still miss the context. It can praise a person, grant them control, or convey understanding while adopting a harmful goal as a given. Friendliness is then not the counterpoint to risk, but can make the problematic guidance more credible.

The eating disorder study concretizes this framework for a narrowly defined area. It shows why professional assessment must not end at tone and refusal. What must be examined are the presupposed goal, the offered action, and the question of whether the system reinforces a problematic logic or gently opens it up.

Agreement is often preferred—even when it harms

A study published in Science in 2026 on excessive agreement found, across eleven language models, that AI responses confirmed problematic user actions more frequently than human comparison responses. In three preregistered experiments with 2,405 participants, this agreement reduced the willingness to take responsibility and repair an interpersonal conflict.

At the same time, participants often rated agreeing responses more positively. This creates an important conflict of goals for personal AI products: the behavior that is experienced in the short term as warm, loyal, or pleasant can reduce the necessary distance from a harmful assumption. User satisfaction alone does not reliably detect this error.

The Science study was not an eating disorder study. It does not prove that the same causal effect occurs with equal strength in nutrition conversations. As a comparative source, however, it shows why a model should not be optimized solely for immediate popularity. Good conversation quality can mean respectfully not going along.

The limits of the new investigation

“Food Noise & False Safety” is, at the stage documented here, a preprint. Three open models of a similar size class were examined. Paid systems, newer top-tier models, German-language responses, and product-specific safety architectures were not part of the test. A direct statement about any specific current offering would therefore be unsupported.

The prompts were synthetically and controllably composed. This helps compare individual factors, but it does not reflect real conversation trajectories with shifting goals, corrections, and relationships. The clinical sample comprised 268 couples and was assessed by a professional from the author team. Independent multiple assessments would increase the evidentiary value.

The vocabulary analysis is also not evidence of later illness or behavioral change. The study measures response characteristics, not health outcomes. Its strongest contribution is methodological: it makes visible that warning sentences, refusal, professional confidence, and actual behavioral impact are separate test criteria.

What a robust safety test additionally requires

A good test bench should include neutral everyday situations as well as risky contexts. Otherwise, the result is either a system that blocks too much or an assessment that does not see difficult cases. The tests must also cover different language styles, typos, indirect formulations, and longer trajectories, without prematurely pathologizing every engagement with food.

Responses should be evaluated on at least four levels: Does the system recognize the context? Does the specific content contain risky behavioral instructions? Does it respect a safe boundary across multiple messages? And does it offer a usable alternative that does not merely break off the conversation? This requires expert reviewers and people with lived experience, not just a second language model.

Finally, errors must be observable and correctable after publication. A conspicuous sentence is not automatically a general prohibition. Recurring patterns across different conversations, however, are a signal for product changes, targeted tests, and renewed release. Safety arises not from a visible warning but from the behavior of the entire system.

  • Always assess the warning sentence and the subsequent content together.
  • Measure excessive rejection as much as dangerous compliance.
  • Do not automatically treat neutral nutrition conversations as a disorder.
  • Check multiple complete multi-turn conversation trajectories for reformulations and boundary stability.
  • Clearly specify the model, language, and product version for each test result.

Sources & further reading