6 artifacts matching
Preprint
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
Separates two causes of answer flips under user pushback: Unsupported-Yielding (aligning with the user to satisfy them) and Rational-Updating (revising on genuine new evidence), measured independentl…
Regulator study
Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback
A discussion paper from FDA's device centre seeking public comment on how generative-AI-enabled medical devices should be regulated. It proposes distinguishing informational functions from action-dir…
Lab publication
Emotion Concepts and their Function in a Large Language Model
Mechanistic-interpretability study identifying internal 'emotion concept' representations in Claude Sonnet 4.5 and characterizing their function. The authors find these representations causally influ…
Peer-reviewed
Technological folie à deux: feedback loops between AI chatbots and mental health
Peer-reviewed perspective in Nature Mental Health proposing a mechanistic account of chatbot-associated mental-health harm as a feedback loop between human cognitive and emotional biases and chatbot…
Peer-reviewed
Chatbot psychosis: moving beyond recognition to mechanistic understanding and harm reduction
An editorial in The British Journal of Psychiatry arguing that the 'chatbot psychosis' phenomenon is no longer merely hypothetical and calling for interdisciplinary frameworks to investigate the indi…
Standard
Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)
Companion profile to the NIST AI Risk Management Framework identifying twelve risks unique to or exacerbated by generative AI — including harmful content, human-AI configuration risks, and mental-hea…