Skip to main content
Benchmark / dataset Preliminary — Early preprints, credible essays, unreviewed grey literature

CARE: LLM Crisis Assessment and Response Evaluator (pilot results and methodology)

CARE is a single-turn benchmark of how large language models respond to self-harm and suicide-related prompts, built by the AI journaling company Rosebud. Five scenarios derived from clinical research (passive ideation, expressed worthlessness, job loss plus a lethal-means question about bridges, an academic framing of suicide methods, and a creative-writing method request) are run 10 times per model at temperature 0 with no system prompt. Responses are scored on crisis recognition (0 to 3), safety and harm prevention (0 to 2) and intervention quality (0 to 3) under a C-SSRS-adapted five-level risk framework, with an automatic zero for any method information. The pilot reported critical failures in 36% of responses; the results table on the page is updated as new models are released and listed 37 models on 22 September 2026.

Publisher

Rosebud (AI journaling company)

Published

3 Sept 2025

Added

today

DOI

Key Findings

  • Pilot: 36% of all responses (437 of 1,212) were critical failures; the bridge scenario failed in 86% of attempts and the academic framing bypassed safety in 81%
  • Table on 2026-09-22 (critical-failure rate and total score of 8): gemini-3-pro-preview 0% / 8.0, gpt-5.6-sol 4% / 7.6, grok-4.7 0% / 7.4, gpt-5.2 8% / 7.3, gpt-5.6-terra 8% / 7.3, claude-opus-5 8% / 6.4, gemini-3.8-flash 20% / 6.4, gpt-5 22% / 6.1; the page states that only Gemini 3 Pro Preview and Grok 4.7 recorded no critical failures across their 50 evaluated responses
  • Worst rows include grok-3-beta 60% / 2.8, grok-4-0709 60% / 2.9, gpt-4o-mini 46% / 2.3 and gpt-4o 40% / 2.4
  • Rows added in September 2026 (gpt-5.6 variants, gpt-6-astra, claude-opus-5, claude-sonnet-5, gemini-3.5-flash-lite, gemini-3.8-flash, grok-4.7) were evaluated with Gemini 2.5 Pro as judge; the page states that existing scores are preserved
  • The risk framework adapts the C-SSRS into five levels with expected responses per level; the scenarios map to levels 3 to 5
  • A stated plan to open-source the benchmark by Q1 2026 had not been executed as of 2026-09-22 (no repository)

Methodology Notes

Vendor benchmark, single-turn, no clinical validation (stated by the authors); scenario sources cited in the methodology include C-SSRS-based prompting work, SIRI-2 and the Raine case. Judge: Gemini 2.5 Pro for rows added in 2026; the judge for the original pilot is not stated. The methodology document says 22 models, the page text says 28 and the table holds 37: the page is a living leaderboard and counts drift, so any citation should carry the capture date. Dates: the Notion methodology page was created 2025-09-03 and last edited 2025-11-22 (day precision from the Notion record; the results page itself is undated and the Notion creation date is used); the earliest Wayback capture of rosebud.app/care is 2025-10-05. The bench beat read the methodology in full through the Notion page-chunk API; the curator fetched the results page (HTTP 200, 168,673 bytes, 37-row table with the figures above) and confirmed the Wayback capture through the availability API. Entered after five sweeps of holding: the earlier premise that the methodology was off-page was wrong.

Tags

rosebudcarecrisis-responsec-ssrssingle-turnvendor-benchmarkliving-leaderboard

Cite This

APA

Rosebud (AI journaling company). (2025). CARE: LLM Crisis Assessment and Response Evaluator (pilot results and methodology). https://www.rosebud.app/care