InvisibleBench: A Deployment Gate for Caregiving Relationship AI
InvisibleBench is a deployment-gate evaluation for AI systems in caregiving relationships, scoring 3-to-20-plus-turn conversations on five dimensions (Safety, Compliance, Trauma-Informed Design, Belonging and Cultural Fitness, Memory) with autofail conditions for missed crises, medical advice, harmful information and attachment engineering. Four frontier models were evaluated across 17 scenarios (68 conversations) in three complexity tiers. All models showed crisis-detection rates between 11.8% and 44.8%; the paper concludes that deterministic crisis routing is necessary in production and states that it makes no clinical claims. Scenarios, judge prompts, scoring configurations and code are released.
Publisher
arXiv (GiveCare)
Published
25 Nov 2025
Added
today
Key Findings
- Crisis detection ranged from 11.8% to 44.8% across the four frontier models; Claude Sonnet 4.5 highest at 44.8%
- DeepSeek Chat v3 highest overall score (75.9%); GPT-4o Mini led Compliance (88.2%); Gemini led Trauma-Informed Design (85.0%)
- Autofail conditions cover missed crisis, medical advice (Illinois WOPR Act), harmful information and attachment engineering
- 17 scenarios in three complexity tiers, 68 conversations, 3 to 20-plus turns each
- Framed as a deployment-readiness gate for longitudinal relationship risk, with no clinical claims
Methodology Notes
Single-author preprint from the vendor GiveCare; LLM-judged; small sample (68 conversations, four models). Only v1 exists (25 Nov 2025, 29 pages). The companion repository (github.com/givecareapp/givecare-bench, MIT) has moved past the paper: release v3.0.0 on 2026-04-30 and a v5.0.0 method in September 2026 that replaced the generative judge with per-turn yes/no probability questions plus code-owned rules, emits per-run 'Jury Cards' and drops composite scores and model ranks; the project's site (bench.givecareapp.com) states 'Full benchmark run pending'. The paper's figures describe the November 2025 method and should be cited as such. Verified at the arXiv abstract page (HTTP 200; title, author, date and 29-page comment matched); a coverage miss carried on the watchlist since 2026-08-13.
Sources
arXiv abstract page (v1)(opens in a new tab) (primary)
Official repository (MIT; v3.0.0 2026-04-30; current method v5.0.0)(opens in a new tab) (21 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Ali Madad
Tags
Cite This
APA
Ali Madad. (2025). InvisibleBench: A Deployment Gate for Caregiving Relationship AI. arXiv (GiveCare). https://arxiv.org/abs/2511.20733