43 artifacts matching
Lab publication
System Card: Claude Sonnet 5.5
System card for Claude Sonnet 5.5, released 28 September 2026, reporting pre-deployment safety, alignment and capability evaluations. Its safeguards chapter reports single-turn and multi-turn results…
NGO report
Artificially Assisted Gun Violence: Chatbot Risks and Preventative Steps for Responsible AI Companies
White paper reviewing firearm-related harms linked to consumer chatbot use, including firearm suicide, attack planning, and illegal acquisition or modification of guns. It audits the published polici…
Lab publication
Claude Opus 5.5 System Card
230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…
Lab publication
Model Card: Grok 4.7
Model card for Grok 4.7, released on 21 September 2026 as xAI's (now styled SpaceXAI) frontier coding and knowledge-work model. Alongside capability benchmarks, the 30-page card reports the company's…
Peer-reviewed
Real-world use and evaluation of a generative AI chatbot for Parkinson's disease information: a prospective observational study
Prospective, conversation-level evaluation of jAImes, a retrieval-augmented Parkinson's disease information chatbot commissioned by Parkinson Stiftung Deutschland and deployed publicly in Germany, ac…
Lab publication
Detecting and countering misuse of AI: September 2026
Anthropic's periodic threat-intelligence report covering December 2025 to August 2026 across seven harm areas. One case study, GTG-15001, documents a China-based app studio that used Claude to build…
Government report
National Commission into the Regulation of AI in Healthcare: Recommendations for a future regulatory framework
Final report of the UK National Commission into the Regulation of AI in Healthcare, established in September 2025 to advise government on a regulatory framework for software and AI-enabled health tec…
Framework
Safe Participation Framework: Opportunity and Safety for the Next Generation in the Age of AI
Microsoft publishes a three-pillar corporate framework for young people's safety across its AI and online products, stating it intends the framework to inform emerging regulatory frameworks. The pill…
Preprint
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini th…
Peer-reviewed
AI chatbot usage may increase mental health help-seeking among college-aged individuals
A randomised source-label experiment testing whether emotional support text is judged differently, and changes help-seeking intentions differently, when it is attributed to ChatGPT rather than to a h…
Lab publication
Continuing To Build Upon Our Safety Priorities
A first-party safety update from Character.AI describing safeguards in operation on its platform as of September 2026. It states that self-harm safeguards consider the surrounding conversation, inclu…
Lab publication
GPT-6 Astra System Card
System card for GPT-6 Astra, published 2026-09-03. Most of the document concerns cyber capabilities at OpenAI's Preparedness 'Critical' threshold, alignment and chain-of-thought monitorability. The p…
Lab publication
System Card: Claude Fable 5.1 & Claude Mythos 5.1
Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…
Peer-reviewed
Privacy assurances and professional-boundary warnings in generative AI mental health chatbots: a randomized vignette experiment on calibrated trust, overreliance risk, and professional help-seeking intentions
2 by 2 randomized vignette experiment with 768 Chinese college students testing whether a privacy-assurance message and a professional-boundary warning in a mental-health chatbot interface change per…
Lab publication
2026 Responsible AI Transparency Report
Microsoft's third annual responsible-AI transparency report, covering July 2025 to June 2026. One of its four 2026 trends is people turning to conversational AI for personal advice, health questions…
Preprint
Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries
An audit of what three free consumer generative-search products (ChatGPT, Perplexity, Google AI Overview) cite when answering mental-health questions. Twenty English questions were run under two prom…
Benchmark / dataset
Announcing Transluce's Mental Health Evaluation (Mental Health Behavior Report)
Independent nonprofit evaluation of how 77 model variants released between May 2024 and July 2026 by OpenAI, Anthropic, Google DeepMind, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI re…
Lab publication
Enabling independent research on how people use Claude
A report on a pilot in which three external research groups designed their own studies and ran them through Anthropic Insights, the company's privacy-preserving aggregate analysis tool, on roughly 25…
Peer-reviewed
Prevalence of Mental Health Discussions in Publicly Available Generative AI Conversations
Cross-sectional research letter analysing 620,699 ChatGPT conversations from the public WildChat-4.8M corpus with an LLM-based classifier that scores topicality, intent, clinical language and affecti…
Preprint
Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
Preregistered experiment in which 1,500 UK adults each held a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was identical across conditions and only the disc…