Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detection), finding that self-harm information crystallizes only in the final 3-7% of layers (93-97% depth), and extracts directional patterns usable for contrastive self-harm detection. Reports that the most accurate probes are not necessarily the most linearly separable, and that Gemma-3-4B encodes these patterns differently from the other tested architectures.
Publisher
arXiv preprint
Published
24 Jul 2026
Added
1 week ago
DOI
—
Key Findings
- Self-harm information becomes linearly decodable only in the final 3-7% of network layers (93-97% depth) across the four tested models
- Directional patterns extracted from probes support contrastive self-harm detection
- The most accurate probes are not necessarily the most linearly separable, and Gemma-3-4B represents self-harm content differently from the other architectures tested
Methodology Notes
arXiv:2607.21988 (cs.CL), v1 submitted 2026-07-24. Author affiliations not stated on the arXiv listing; recorded by author names. Linear probes across all layers of four models; two self-harm datasets (X-Sensitive, SH-Detection). Framed as supporting self-harm detection systems and LLM governance.
Sources
arXiv abstract page (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Luis Espinosa-Anke, Carla Perez-Almendros
Tags
Cite This
APA
Luis Espinosa-Anke, Carla Perez-Almendros (2026). Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study. arXiv preprint. https://arxiv.org/abs/2607.21988