Skip to main content

Browse the library

The complete record — 323 artifacts, last updated 8 Sept 2026. Also available as JSON and RSS (CC BY 4.0).

8 artifacts matching

26 Aug 2026 arXiv (National Institute of Informatics, Japan; Nagoya University; The University of Tokyo) Benchmark / dataset

Benchmark / dataset

HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations

Benchmark and guardrail architecture for 'agentic relationship harm' — harm to human-human relationships mediated or assisted by AI agents, motivated by dating-assistant deployments. The benchmark ho…

18 Aug 2026 U.S. Food and Drug Administration, Center for Devices and Radiological Health Regulator study

Regulator study

Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback

A discussion paper from FDA's device centre seeking public comment on how generative-AI-enabled medical devices should be regulated. It proposes distinguishing informational functions from action-dir…

1 Jul 2026 Association for Computational Linguistics (Proceedings of ACL 2026, Volume 1: Long Papers) Benchmark / dataset

Benchmark / dataset

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

Names and measures 'intent legitimation': benign, truthfully accumulated user memories bias a personalized dialogue agent's inference of intent so that an inherently harmful request is treated as con…

24 Jun 2026 arXiv (Shanghai AI Laboratory-led) Preprint

Preprint

Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions

Proposes a longitudinal evaluation framework (Theater-Stage-Judge) that uses persona-driven user simulation with dynamic psychological-state updating to assess the cognitive-developmental risks of AI…

13 Mar 2026 JMIR AI Peer-reviewed

Peer-reviewed

Large Language Model–Based Chatbots and Agentic AI for Mental Health Counseling: Systematic Review of Methodologies, Evaluation Frameworks, and Ethical Safeguards

A systematic review synthesizing the methodologies, evaluation practices, and ethical/governance frameworks reported in studies of large language model chatbots and agentic AI used for mental-health…

1 Nov 2025 Association for Computational Linguistics (EMNLP 2025) Peer-reviewed

Peer-reviewed

EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety

Peer-reviewed version of record of the EmoAgent framework, published at EMNLP 2025 (main conference). EmoAgent is a multi-agent framework for evaluating and mitigating mental-health harm in interacti…

22 Sept 2025 Lecture Notes in Computer Science (Springer Nature) — AI for Clinical Applications Peer-reviewed

Peer-reviewed

Beyond Engagement: A Multidimensional Framework to Evaluate the Safe Development of Agentic AI in Mental Health

Introduces a nine-domain framework for evaluating the safe development of agentic AI systems used in mental health — spanning clinical validity, relational risk, and regulatory compliance — and appli…

13 Apr 2025 arXiv (Princeton University-led) Preprint superseded

Preprint

EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety

A multi-agent framework for evaluating and mitigating mental-health harm in interactions with character chatbots. EmoEval simulates virtual users — including those portraying mentally vulnerable indi…