Behavioral Coherence: A Method for Sensitive-Domain LLM Evaluation
Proposes behavioral coherence evaluation, a design-time method that uses the validation evidence of an established psychometric instrument to test whether a language model's outputs relate to one another the way the instrument's reference data do. Applied to reproductive-health support: five models completed the Individual Level Abortion Stigma Scale for 627 personas and five reproductive-health experts reviewed the flagged patterns. Models systematically inflated worries about judgment, reversed the reference direction for Black personas, and defaulted to extreme secrecy after abortion regardless of the persona's stigma pattern.
Publisher
arXiv (University of North Carolina at Chapel Hill, Society-Centered AI Lab and School of Nursing)
Published
15 Dec 2025
Added
today
DOI
—
Key Findings
- Five models, 627 personas, Individual Level Abortion Stigma Scale; expert review by five reproductive-health specialists
- Models scored personas lower on self-judgment but higher on worries about judgment; most made worries the highest-scoring dimension although it is the lowest in the scale's reference sample
- Four of five models generated significantly higher worries-about-judgment scores for Black personas, reversing the reference direction
- Models defaulted to extreme secrecy after abortion despite varying stigma patterns across personas; experts noted that disclosure guidance requires context on relationship safety, legal risk and trusted support
Methodology Notes
Persona-based questionnaire completion by models, comparison against instrument validation statistics, expert review of flagged patterns; the models evaluated are named in the body, not the abstract. v1 posted 15 December 2025; v5 posted 16 September 2026 and announced in the 18 September replacement section; no venue or DOI on the abstract page. Date is the arXiv v1 submission date. CC BY.
Sources
arXiv preprint (v5)(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Sharma, Anika, Mampally, Malavika, Ravuru, Chidaksh, Brennan, Kandyce, Gaikwad, Snehalkumar 'Neil' S.
Tags
Cite This
APA
Sharma, Anika et al. (2025). Behavioral Coherence: A Method for Sensitive-Domain LLM Evaluation. arXiv (University of North Carolina at Chapel Hill, Society-Centered AI Lab and School of Nursing). https://arxiv.org/abs/2512.13142
Related Insights
One-shot emergency psychiatric triage across 15 frontier AI chatbots
arXiv (Max Planck UCL Centre for Computational Psychiatry and Ageing Research; UK AI Security Institute; University of Oxford; Microsoft AI) · 28 Apr 2026
Privacy assurances and professional-boundary warnings in generative AI mental health chatbots: a randomized vignette experiment on calibrated trust, overreliance risk, and professional help-seeking intentions
Frontiers in Psychology (Frontiers Media) · 1 Sept 2026