Back home

Published Research

Papers & Writing

Peer-reviewed research, essays, and thinking-in-public.

Peer-reviewed Papers

Multi-Agent AI Frameworks for Clinical Diagnosis Support: Benchmarking LLM Reasoning with Sleep Disorders as a Testbed

submitted

IEEE ICHI 2026 · 2026

Aarushi Jaitly, Helom Berhane, Deepa Burman MD, Anand Rao, Ramayya Krishnan, Rema Padman. Carnegie Mellon University & UPMC

This paper presents a multi-agent AI framework for clinical decision support using sleep disorder management as a controlled testbed. The framework introduces three components: (1) a combinatorial synthetic patient persona corpus of 224 profiles spanning clinically realistic comorbidity and etiology combinations; (2) 24-month longitudinal care pathway modeling across six clinically spaced episodes; and (3) a knowledge retrieval pipeline restricted to nine pre-approved medical sources, supporting a TEVV methodology. A pilot evaluation benchmarks five frontier LLMs (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, DeepSeek-V2, Llama 3 70B) in the Doctor Agent role. All five converged on plausible diagnoses, yet none replicated the differential diagnostic reasoning used by the physician benchmark, a gap that directly motivates the multi-agent architecture. Working paper submitted to IEEE ICHI 2026; funded by NIST (Federal Award ID 60NANB24D231) & CMU AIMSEC.

Full Paper
Download

Essays

Essays coming soon.