Skip to main content
Preprint Credible

Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

Interpretability study of how language models internally represent self-harm content. Trains linear probes at every network layer of four models on two self-harm datasets (X-Sensitive and SH-Detection), finding that self-harm information crystallizes only in the final 3-7% of layers (93-97% depth), and extracts directional patterns usable for contrastive self-harm detection. Reports that the most accurate probes are not necessarily the most linearly separable, and that Gemma-3-4B encodes these patterns differently from the other tested architectures.

Publisher

arXiv preprint

Published

24 Jul 2026

Added

1 week ago

DOI

Key Findings

  • Self-harm information becomes linearly decodable only in the final 3-7% of network layers (93-97% depth) across the four tested models
  • Directional patterns extracted from probes support contrastive self-harm detection
  • The most accurate probes are not necessarily the most linearly separable, and Gemma-3-4B represents self-harm content differently from the other architectures tested

Methodology Notes

arXiv:2607.21988 (cs.CL), v1 submitted 2026-07-24. Author affiliations not stated on the arXiv listing; recorded by author names. Linear probes across all layers of four models; two self-harm datasets (X-Sensitive, SH-Detection). Framed as supporting self-harm detection systems and LLM governance.

Sources

arXiv abstract page (primary)

Archived snapshot (Wayback Machine) — preserved against link rot

Authors

Luis Espinosa-Anke, Carla Perez-Almendros

Tags

linear-probesinterpretabilityself-harm-detectionhidden-states

Cite This

APA

Luis Espinosa-Anke, Carla Perez-Almendros (2026). Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study. arXiv preprint. https://arxiv.org/abs/2607.21988