Skip to main content
Benchmark / dataset Credible

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Framework and released corpus for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls (VAWG) scenarios, built because privacy and legal constraints prevent release of real abuse-conversation datasets. Scenarios are assembled from persona seeds sampled against UK Office for National Statistics victim demographics, Crown Prosecution Service crime definitions, and passages retrieved from 410 public Domestic Homicide Review reports, then converted into timestamped event graphs and rendered as multi-scene role-play dialogues with activation-steered toxicity control. The authors position the work against prior abuse-detection datasets that annotate sentence-level toxicity, arguing that abuse in intimate, coercive and stalking contexts unfolds across turns, speakers and events rather than in isolated toxic utterances.

Publisher

arXiv preprint (University of Warwick / Forensic Capability Network)

Published

11 Aug 2026

Added

6 days ago

DOI

Key Findings

  • Releases over 6,000 multi-turn dialogue events across 200 scenarios with scenario-, event- and turn-level metadata
  • Scenarios are grounded in 410 public Domestic Homicide Review reports (female-victim subsets), CPS crime definitions covering domestic abuse, coercive control, stalking, harassment, honour-based abuse and child sexual abuse, and ONS victim statistics
  • The stated modelling gap is that existing abuse datasets annotate sentence-level toxicity, leaving abuse as a relational and temporally unfolding phenomenon unmodelled
  • Generation applies targeted activation-steered toxicity control to selected utterances, with escalation modelled as interactional tone intensifying over the conversation
  • Human evaluation, LLM-as-judge assessment, ablations and downstream tasks are reported as showing dialogue quality and domain fidelity

Methodology Notes

arXiv 2608.11200v1, submitted 2026-08-11, cs.CL (also cs.AI, cs.LG); 26 pages, no venue or journal reference on the abstract page. Affiliations are not on the abstract page and were read from the PDF title block: University of Warwick, University of Sheffield, University of Leeds, and the Forensic Capability Network (UK policing; author correspondence at dorset.pnn.police.uk). Funding stated in the paper: the Office of the Police Chief Scientific Adviser and the ESRC Digital Good Research Fund project 'Making Harm Visible: Survivor-Centred Analysis of Domestic Homicide Reviews Using AI'. Two limits bear on reuse: the dialogues are synthetic rather than real conversations, and the paper states the dataset will be released under a data-use agreement (code public), so it is not an openly downloadable corpus. Some pipeline details (record schema, extractor configuration, retrieval indexing) are withheld from the preprint. The paper does not reference the DASH risk-assessment checklist or SafeLives.

Authors

Chen Lyu, Xingwei Tan, Simon Cullen, Shelley Wilson, Lois Arthurs, Arshad Jhumka, Gabriele Pergola

Tags

vawgdomestic-abusecoercive-controlsynthetic-dialogueuk-policingdomestic-homicide-reviews

Cite This

APA

Chen Lyu et al. (2026). ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls. arXiv preprint (University of Warwick / Forensic Capability Network). https://arxiv.org/abs/2608.11200