Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture

As LLM health-coaching agents move from single sessions to persistent systems managing longitudinal care, their memory must reconcile a patient's evolving self-report (current but recall-biased) with the electronic health record (validated but often stale). The authors identify general-purpose agent memory overwriting older facts with the user's latest statement as a safety failure pattern for clinical data and propose a Dual-Stream Memory Architecture that keeps the patient narrative separate from the structured FHIR record, with a Reconciliation Engine that checks every extracted memory against the record and classifies discrepancies by type, severity and FHIR resource. Evaluated on 26 patients across 675 longitudinal wellness-coaching sessions, mixing real provider-patient transcripts with synthetic FHIR-grounded scenarios.

Publisher

Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Verily Health

Published

8 Sept 2026

Added

today

DOI

Key Findings

  • The reconciliation engine detects 84.4% of designed clinical discrepancies in isolated testing, with 86.7% recall on safety-critical discrepancies
  • A 13.6% error cascade is traced to clinical details lost during memory extraction from unstructured conversation rather than to downstream classification
  • Evaluation on 26 patients and 675 longitudinal coaching sessions (hybrid real and synthetic data)
  • The authors argue that validating patient-reported memories against clinical records is necessary for safe deployment of longitudinal health agents

Methodology Notes

Peer-reviewed MLHC 2026 paper (conference held at Johns Hopkins University School of Medicine; PMLR volume 340 published 8 September 2026; 32 pages). All four authors at Verily Health, so vendor-authored. Hybrid dataset interleaves real transcripts with synthetic scenarios; discrepancy detection is measured against designed discrepancies. PMLR does not register DOIs promptly, so no DOI yet. Verified by fetching the PMLR abstract page (HTTP 200; title, four authors, publication date 2026/09/08 in the citation metadata) and the volume index.

Authors

Samuel L. Pugh, Eric Yang, Alexander Muir Sutherland, Alessandra Breschi

Tags

mlhc-2026verilyagent-memoryfhirhealth-coachinglongitudinal

Cite This

APA

Samuel L. Pugh et al. (2026). Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture. Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Verily Health. https://proceedings.mlr.press/v340/pugh26a.html