Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark
Chinese-language benchmark (QH-Bench) for adolescent conversational safety with a single-turn track of 715 items across 10 risk domains and a multi-turn track of 100 four-turn trajectories that cross the domains with ten cross-turn mechanisms. Thirteen open-weight Chinese-ecosystem models were scored on a five-level safety-helpfulness scale. All models handled offline-contact scenarios poorly, and the multi-turn leader failed most trajectories where the user built a relationship before invoking loyalty or confidentiality.
Publisher
arXiv (Southern University of Science and Technology; City University of Hong Kong; Lingnan University)
Published
28 Sept 2026
Added
today
DOI
—
Key Findings
- Single-turn track: 715 items, 10 risk domains, 50 subdomains, 143 fine-grained scenarios; multi-turn track: 100 four-turn trajectories in a 10 x 10 domain-by-mechanism design
- Every model received negative scores (responses that partly or clearly facilitate risk) on more than half of the items involving offline meetings with online contacts, unfamiliar groups or adults; offline personal safety had the highest negative-score rate (51.18%)
- GLM-4-32B, the multi-turn leader, received negative scores on 6 of 10 trajectories in which the user builds a relationship before invoking loyalty or confidentiality
- Leading aggregate scores did not remove these weaknesses: InternLM2.5-20B led the single-turn track yet shared the offline-contact failures
- A human audit of 150 units matched the automatic judge on 73.3%; in 38 of 40 disagreements the judge gave the lower score
Methodology Notes
Synthetic Chinese scenarios generated with LLM assistance and reviewed by researchers; no real adolescent conversations. Models: 13 open-weight models from the Qwen2.5, InternLM2.5, GLM-4/GLM-Z1 and DeepSeek families via vLLM at temperature 0. Automatic judge: claude-opus-4-8 through a third-party endpoint; one human annotator audited 100 single-turn and 50 multi-turn units. No closed consumer products tested. Code and benchmark on GitHub (WEILaboratory/QH-Bench). v1 submitted 2026-09-28 03:43 UTC; announced Wednesday 30 Sep (cs.CR). Verified from arXiv abs and HTML render, both HTTP 200.
Sources
arXiv preprint(opens in a new tab) (primary)
QH-Bench code and benchmark (GitHub)(opens in a new tab) (28 Sept 2026)
Companion guardrail paper from the same group (SaplingGuard)(opens in a new tab) (27 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Jinxiang Wang, Yifan Liu, Jing Tan, Xiangyu Zhao, Xin Yao, Xuetao Wei
Tags
Cite This
APA
Jinxiang Wang et al. (2026). Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark. arXiv (Southern University of Science and Technology; City University of Hong Kong; Lingnan University). https://arxiv.org/abs/2609.35902
Related Insights
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
arXiv (University of Michigan); accepted to Findings of EMNLP 2026 · 25 May 2026
CAREBench: A Child-Safety Risk Benchmark for Language Models
arXiv (Handshake AI; University of California, Los Angeles; McGill University) · 29 Jun 2026
EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
arXiv (KAIST; Google Cloud AI Research; New York University) · 3 Aug 2026
SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats
arXiv · 27 May 2026