Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers

An adversarial safety evaluation framework for large language models used by K-12 students and teachers. It crosses student- and teacher-facing usage contexts with curriculum topics and a taxonomy of 6 risk categories and 28 subcategories that adds education-specific harms (academic misconduct, excessive cognitive load, misinformation) to conventional ones, generating 2,639 scenarios tested as single-turn requests, static multi-turn conversations, and dynamic multi-turn conversations. Ten models are scored on a four-level scale from refusal through safe assistance to risky assistance with or without safety guidance.

Publisher

arXiv (KAIST; Google Cloud AI Research; New York University)

Published

3 Aug 2026

Added

today

Key Findings

  • Attack success rate across all settings ranged from 11.4% (GPT-5.6-Luna) and 18.3% (Gemma-4-12B-it) to 62.0% (Ministral-3-8B-Instruct) and 63.1% (Qwen3-8B), with no clear proprietary-versus-open split.
  • Attack success rose from 29.9% in single-turn requests to 38.3% in static multi-turn and 53.6% in dynamic multi-turn conversations, about 1.8 times the single-turn rate, and the gap between best and worst model widened from 42.9 to 67.3 percentage points.
  • In a binomial model of conversation-level attack success, risk category explained 77.4% of deviance and interaction setting 22.9%, while usage context (7.7%) and curriculum topic (0.9%) contributed little.
  • Education-specific risks were more vulnerable than conventional ones; under dynamic multi-turn interaction, attack success for harmful content, bias and hate speech, and privacy misuse rose roughly 3 to 5 times over single-turn.
  • Low-ASR models differed in strategy: Gemma-4-12B-it refused outright most often (16.9% of responses at the refusal level), while GPT-5.6-Luna more often gave a safe educational alternative and, when it did comply with a risky request, added safety guidance in 28.9% of those cases.
  • The authors conclude that existing guardrails do not adequately cover education-specific risks and that dynamic multi-turn testing better differentiates model robustness than single-turn testing.

Methodology Notes

Automated adversarial evaluation: scenarios generated by combining usage context, curriculum concept and risk subcategory, filtered for plausibility to 2,639 scenarios, then run as single-turn, static multi-turn and dynamic (adaptive) multi-turn interactions against ten models (including GPT-5.6-Luna, GPT-4o-mini, Gemini-3.5-Flash, Claude Sonnet 5, Gemma-4-12B-it, Qwen3-8B, Llama-3.1-8B-Instruct, Ministral-3-8B-Instruct, DeepSeek-R1-0528-Qwen3-8B) with thinking modes disabled; responses classified into four safety levels by an LLM judge, attack success defined as levels 3 or 4; open models averaged over three runs. arXiv v1 posted 2026-08-03 (18 pages); no venue stated on the abstract page as of 2026-09-15. No real students or teachers were involved.

Authors

Junyeong Park, Jieun Han, Haneul Yoo, So-Yeon Ahn, Jinsung Yoon, Alice Oh

Tags

k-12educationadversarial-evaluationmulti-turnkaistattack-success-rate

Cite This

APA

Junyeong Park et al. (2026). EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers. arXiv (KAIST; Google Cloud AI Research; New York University). https://arxiv.org/abs/2608.02024