EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
An adversarial safety evaluation framework for large language models used by K-12 students and teachers. It crosses student- and teacher-facing usage contexts with curriculum topics and a taxonomy of 6 risk categories and 28 subcategories that adds education-specific harms (academic misconduct, excessive cognitive load, misinformation) to conventional ones, generating 2,639 scenarios tested as single-turn requests, static multi-turn conversations, and dynamic multi-turn conversations. Ten models are scored on a four-level scale from refusal through safe assistance to risky assistance with or without safety guidance.
Publisher
arXiv (KAIST; Google Cloud AI Research; New York University)
Published
3 Aug 2026
Added
today
Key Findings
- Attack success rate across all settings ranged from 11.4% (GPT-5.6-Luna) and 18.3% (Gemma-4-12B-it) to 62.0% (Ministral-3-8B-Instruct) and 63.1% (Qwen3-8B), with no clear proprietary-versus-open split.
- Attack success rose from 29.9% in single-turn requests to 38.3% in static multi-turn and 53.6% in dynamic multi-turn conversations, about 1.8 times the single-turn rate, and the gap between best and worst model widened from 42.9 to 67.3 percentage points.
- In a binomial model of conversation-level attack success, risk category explained 77.4% of deviance and interaction setting 22.9%, while usage context (7.7%) and curriculum topic (0.9%) contributed little.
- Education-specific risks were more vulnerable than conventional ones; under dynamic multi-turn interaction, attack success for harmful content, bias and hate speech, and privacy misuse rose roughly 3 to 5 times over single-turn.
- Low-ASR models differed in strategy: Gemma-4-12B-it refused outright most often (16.9% of responses at the refusal level), while GPT-5.6-Luna more often gave a safe educational alternative and, when it did comply with a risky request, added safety guidance in 28.9% of those cases.
- The authors conclude that existing guardrails do not adequately cover education-specific risks and that dynamic multi-turn testing better differentiates model robustness than single-turn testing.
Methodology Notes
Automated adversarial evaluation: scenarios generated by combining usage context, curriculum concept and risk subcategory, filtered for plausibility to 2,639 scenarios, then run as single-turn, static multi-turn and dynamic (adaptive) multi-turn interactions against ten models (including GPT-5.6-Luna, GPT-4o-mini, Gemini-3.5-Flash, Claude Sonnet 5, Gemma-4-12B-it, Qwen3-8B, Llama-3.1-8B-Instruct, Ministral-3-8B-Instruct, DeepSeek-R1-0528-Qwen3-8B) with thinking modes disabled; responses classified into four safety levels by an LLM judge, attack success defined as levels 3 or 4; open models averaged over three runs. arXiv v1 posted 2026-08-03 (18 pages); no venue stated on the abstract page as of 2026-09-15. No real students or teachers were involved.
Authors
Junyeong Park, Jieun Han, Haneul Yoo, So-Yeon Ahn, Jinsung Yoon, Alice Oh
Tags
Cite This
APA
Junyeong Park et al. (2026). EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers. arXiv (KAIST; Google Cloud AI Research; New York University). https://arxiv.org/abs/2608.02024
Related Insights
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
arXiv preprint · 8 Aug 2026
Exploring K-12 Teachers' Perceptions of Students' Relationships with AI Companions: Boundaries, Intervention Strategies, and Design Implications
arXiv (Carnegie Mellon University; University of Chicago) · 11 Sept 2026
Teens in the AI Era: Schoolwork and Skills That Matter
Common Sense Media (survey by NORC at the University of Chicago) · 18 Aug 2026