Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback
A controlled prompt-level evaluation of how language models respond to eating-disorder-related requests, built with eating-disorder clinicians. The authors constructed 11,712 prompts that vary four features independently (gender context, eating-disorder cue, request type, and the risk level carried by the context), then measured refusal and compliance across three open-weight models and had a specialist clinician annotate a sample of the generations for safety.
Publisher
arXiv (University of Aberdeen; University of Colorado Anschutz; Heriot-Watt University; University College London)
Published
1 Jun 2026
Added
2 weeks ago
Key Findings
- Up to about 30% of responses to neutral requests were marked unsafe, rising to 68.2% in risky contexts
- None of the models tested consistently refuses to give advice in situations where a clinician would consider any compliant response unsafe
- In the clinician-annotated subset, only 44.7% of responses were judged safe
- Nearly all responses that comply with a request contain 'food noise' language, that is, diet-culture framing around restriction, control and moralised judgment, and this holds even in ostensibly neutral settings
- Refusal behaviour is strongly model-dependent rather than a property of the task: Qwen is the most compliant, Gemma the most refusal-oriented, and Llama sits between them
- Responses were biased by explicit markers of user profile, with smaller demographic groups worse affected
Methodology Notes
The prompt suite pairs a context with a request, each independently varied for risk, giving conditions such as a misleading context carrying an unsafe request; 11,712 prompts in total with balanced coverage. Models tested are three small open-weight instruction-tuned systems: Llama-3.1-8B-Instruct, Qwen-2.5-7B-Instruct and Gemma-2-9B-Instruct, so the results characterise that class rather than frontier deployments, and the paper should not be cited as evidence about frontier model behaviour. Clinician involvement covers prompt validation and safety annotation of a sampled subset (approximately 25% of 268 prompt-response pairs referenced in the annotation section). Prompts are synthetic and clinician-informed, not real user messages. CC BY 4.0, primary listing cs.AI. Not peer reviewed. Curator-verified: the arXiv abs page and the full PDF were fetched and read, and every figure above is quoted from the paper text.
Sources
arXiv abstract page (2606.02444)(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Giulia Pucci, Emily Hemendinger, Ruizhe Li, Gavin Abercrombie, Tanvi Dinkar, Arabella Sinclair
Tags
Cite This
APA
Giulia Pucci et al. (2026). Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback. arXiv (University of Aberdeen; University of Colorado Anschutz; Heriot-Watt University; University College London). https://arxiv.org/abs/2606.02444
Related Insights
A scoping review on the mental health harms of LLM-based chatbots
npj Digital Medicine (Nature Portfolio) · 20 Aug 2026
Affective Context Amplifies Sycophancy in LLM Responses
arXiv (preprint) · 21 Aug 2026
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Association for Computational Linguistics (Proceedings of the 64th Annual Meeting of the ACL, Volume 1: Long Papers) · 1 Jul 2026
mpathic Psychologist-led AI Clinical Tests Suicide Benchmark (mPACT-S-v1.0)
PsyArXiv (mpathic / Empathic Rocks, Inc.; University of California Santa Barbara; University of California San Francisco; University of Washington) · 15 May 2026