4 artifacts matching
Emotion Concepts and their Function in a Large Language Model
Mechanistic-interpretability study identifying internal 'emotion concept' representations in Claude Sonnet 4.5 and characterizing their function. The authors find these representations causally influ…
Technological folie à deux: feedback loops between AI chatbots and mental health
Peer-reviewed perspective in Nature Mental Health proposing a mechanistic account of chatbot-associated mental-health harm as a feedback loop between human cognitive and emotional biases and chatbot…
Chatbot psychosis: moving beyond recognition to mechanistic understanding and harm reduction
An editorial in The British Journal of Psychiatry arguing that the 'chatbot psychosis' phenomenon is no longer merely hypothetical and calling for interdisciplinary frameworks to investigate the indi…
Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)
Companion profile to the NIST AI Risk Management Framework identifying twelve risks unique to or exacerbated by generative AI — including harmful content, human-AI configuration risks, and mental-hea…