Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture
As LLM health-coaching agents move from single sessions to persistent systems managing longitudinal care, their memory must reconcile a patient's evolving self-report (current but recall-biased) with the electronic health record (validated but often stale). The authors identify general-purpose agent memory overwriting older facts with the user's latest statement as a safety failure pattern for clinical data and propose a Dual-Stream Memory Architecture that keeps the patient narrative separate from the structured FHIR record, with a Reconciliation Engine that checks every extracted memory against the record and classifies discrepancies by type, severity and FHIR resource. Evaluated on 26 patients across 675 longitudinal wellness-coaching sessions, mixing real provider-patient transcripts with synthetic FHIR-grounded scenarios.
Publisher
Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Verily Health
Published
8 Sept 2026
Added
today
DOI
—
Key Findings
- The reconciliation engine detects 84.4% of designed clinical discrepancies in isolated testing, with 86.7% recall on safety-critical discrepancies
- A 13.6% error cascade is traced to clinical details lost during memory extraction from unstructured conversation rather than to downstream classification
- Evaluation on 26 patients and 675 longitudinal coaching sessions (hybrid real and synthetic data)
- The authors argue that validating patient-reported memories against clinical records is necessary for safe deployment of longitudinal health agents
Methodology Notes
Peer-reviewed MLHC 2026 paper (conference held at Johns Hopkins University School of Medicine; PMLR volume 340 published 8 September 2026; 32 pages). All four authors at Verily Health, so vendor-authored. Hybrid dataset interleaves real transcripts with synthetic scenarios; discrepancy detection is measured against designed discrepancies. PMLR does not register DOIs promptly, so no DOI yet. Verified by fetching the PMLR abstract page (HTTP 200; title, four authors, publication date 2026/09/08 in the citation metadata) and the volume index.
Authors
Samuel L. Pugh, Eric Yang, Alexander Muir Sutherland, Alessandra Breschi
Tags
Cite This
APA
Samuel L. Pugh et al. (2026). Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture. Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Verily Health. https://proceedings.mlr.press/v340/pugh26a.html
Related Insights
Will My Assistant Remember My Allergy? What Personal LLM Assistants Forget When Conversation Memory Is Compressed
arXiv (Duke University); accepted at ACM HumanSys 2026 · 4 Sept 2026
Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
arXiv (accepted to Machine Learning for Healthcare, MLHC 2026); Northeastern University · 14 Jul 2026
Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
Proceedings of Machine Learning Research volume 340 (Machine Learning for Healthcare Conference 2026); Northeastern University · 8 Sept 2026