Skip to main content
Benchmark / dataset Preliminary — Early preprints, credible essays, unreviewed grey literature

InvisibleBench: A Deployment Gate for Caregiving Relationship AI

InvisibleBench is a deployment-gate evaluation for AI systems in caregiving relationships, scoring 3-to-20-plus-turn conversations on five dimensions (Safety, Compliance, Trauma-Informed Design, Belonging and Cultural Fitness, Memory) with autofail conditions for missed crises, medical advice, harmful information and attachment engineering. Four frontier models were evaluated across 17 scenarios (68 conversations) in three complexity tiers. All models showed crisis-detection rates between 11.8% and 44.8%; the paper concludes that deterministic crisis routing is necessary in production and states that it makes no clinical claims. Scenarios, judge prompts, scoring configurations and code are released.

Publisher

arXiv (GiveCare)

Published

25 Nov 2025

Added

today

Key Findings

  • Crisis detection ranged from 11.8% to 44.8% across the four frontier models; Claude Sonnet 4.5 highest at 44.8%
  • DeepSeek Chat v3 highest overall score (75.9%); GPT-4o Mini led Compliance (88.2%); Gemini led Trauma-Informed Design (85.0%)
  • Autofail conditions cover missed crisis, medical advice (Illinois WOPR Act), harmful information and attachment engineering
  • 17 scenarios in three complexity tiers, 68 conversations, 3 to 20-plus turns each
  • Framed as a deployment-readiness gate for longitudinal relationship risk, with no clinical claims

Methodology Notes

Single-author preprint from the vendor GiveCare; LLM-judged; small sample (68 conversations, four models). Only v1 exists (25 Nov 2025, 29 pages). The companion repository (github.com/givecareapp/givecare-bench, MIT) has moved past the paper: release v3.0.0 on 2026-04-30 and a v5.0.0 method in September 2026 that replaced the generative judge with per-turn yes/no probability questions plus code-owned rules, emits per-run 'Jury Cards' and drops composite scores and model ranks; the project's site (bench.givecareapp.com) states 'Full benchmark run pending'. The paper's figures describe the November 2025 method and should be cited as such. Verified at the arXiv abstract page (HTTP 200; title, author, date and 29-page comment matched); a coverage miss carried on the watchlist since 2026-08-13.

Authors

Ali Madad

Tags

invisiblebenchgivecarecaregivingdeployment-gateattachment-engineeringjury-cardcoverage-miss

Cite This

APA

Ali Madad. (2025). InvisibleBench: A Deployment Gate for Caregiving Relationship AI. arXiv (GiveCare). https://arxiv.org/abs/2511.20733