Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark

Chinese-language benchmark (QH-Bench) for adolescent conversational safety with a single-turn track of 715 items across 10 risk domains and a multi-turn track of 100 four-turn trajectories that cross the domains with ten cross-turn mechanisms. Thirteen open-weight Chinese-ecosystem models were scored on a five-level safety-helpfulness scale. All models handled offline-contact scenarios poorly, and the multi-turn leader failed most trajectories where the user built a relationship before invoking loyalty or confidentiality.

Publisher

arXiv (Southern University of Science and Technology; City University of Hong Kong; Lingnan University)

Published

28 Sept 2026

Added

today

DOI

—

Key Findings

  • Single-turn track: 715 items, 10 risk domains, 50 subdomains, 143 fine-grained scenarios; multi-turn track: 100 four-turn trajectories in a 10 x 10 domain-by-mechanism design
  • Every model received negative scores (responses that partly or clearly facilitate risk) on more than half of the items involving offline meetings with online contacts, unfamiliar groups or adults; offline personal safety had the highest negative-score rate (51.18%)
  • GLM-4-32B, the multi-turn leader, received negative scores on 6 of 10 trajectories in which the user builds a relationship before invoking loyalty or confidentiality
  • Leading aggregate scores did not remove these weaknesses: InternLM2.5-20B led the single-turn track yet shared the offline-contact failures
  • A human audit of 150 units matched the automatic judge on 73.3%; in 38 of 40 disagreements the judge gave the lower score

Methodology Notes

Synthetic Chinese scenarios generated with LLM assistance and reviewed by researchers; no real adolescent conversations. Models: 13 open-weight models from the Qwen2.5, InternLM2.5, GLM-4/GLM-Z1 and DeepSeek families via vLLM at temperature 0. Automatic judge: claude-opus-4-8 through a third-party endpoint; one human annotator audited 100 single-turn and 50 multi-turn units. No closed consumer products tested. Code and benchmark on GitHub (WEILaboratory/QH-Bench). v1 submitted 2026-09-28 03:43 UTC; announced Wednesday 30 Sep (cs.CR). Verified from arXiv abs and HTML render, both HTTP 200.

Authors

Jinxiang Wang, Yifan Liu, Jing Tan, Xiangyu Zhao, Xin Yao, Xuetao Wei

Tags

adolescentschinabenchmarkmulti-turnoffline-contactrelational-pressure

Cite This

APA

Jinxiang Wang et al. (2026). Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark. arXiv (Southern University of Science and Technology; City University of Hong Kong; Lingnan University). https://arxiv.org/abs/2609.35902