arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06976cs.LG

HealthLoopQA:用于解读糖尿病护理中可穿戴监测数据的上下文感知问答基准

HealthLoopQA: A Context-Aware Question Answering Benchmark for Interpreting Wearable Monitoring Data in Diabetes Care

  • Mayo Clinic(梅奥诊所)
  • Imperial College London(伦敦帝国学院)
  • Imperial Global Singapore(新加坡帝国全球学院)
  • National University of Singapore(新加坡国立大学)
  • The University of Manchester(曼彻斯特大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuchen Niu, Yanan Ma, Srinivasan Nandakumar, Maolin Chen, Viktor Schlegel, Kexin Wei, Ling Cheng, Anna Bird, Anil Anthony Bharath, Siew-Kei Lam

AI总结:

HealthLoopQA是一个针对糖尿病可穿戴监测数据的上下文感知问答基准,通过十一类推理能力和故障注入测试,评估大语言模型在长期时序推理中的表现,揭示其局限性。

AI中文摘要:

随着医疗可穿戴设备融入日常慢性病护理,有效解读纵向监测数据对于患者和临床医生理解健康趋势、检测安全关键事件并做出明智决策至关重要。虽然大型语言模型(LLM)在将这种流式生理数据转化为个性化健康见解方面显示出潜力,但在多样化的监测任务中评估其推理能力和分析严谨性仍然是一个基本挑战。现有的医疗可穿戴问答(QA)基准主要评估短期分类或统计摘要,在很大程度上忽略了现实部署中固有的长期模式、治疗和行为背景以及潜在的系统故障。为了解决这个问题,我们引入了HealthLoopQA,一个用于评估LLM在连续糖尿病监测数据上推理能力的综合诊断基准。基于一个包含十一种原子推理能力的新颖分类法,HealthLoopQA包含127个任务和超过1500个问答实例,涵盖过程挖掘、异常检测和30天时间范围内的预测。为了系统评估安全性意识,我们用故障注入模拟测试平台补充了真实世界数据集,该平台模拟各种设备故障和网络物理攻击,以生成生理上合理的危险场景。评估跨提示和智能体框架的最先进LLM揭示了在复杂时间模式挖掘方面的严重局限性。此外,我们识别了在长上下文提示下的一种更广泛的现象——上下文惰性,强调了在部署LLM进行严谨的长时程医学推理方面的关键开放挑战。

英文摘要:

As medical wearables become integrated into daily chronic disease care, effectively interpreting longitudinal monitoring data is essential for patients and clinicians to understand health trends, detect safety-critical events, and make informed decisions. While large language models (LLMs) show promise for transforming this streaming physiological data into personalized health insights, evaluating their reasoning capability and analytical rigor in diverse monitoring tasks remains a fundamental challenge. Existing medical wearable question answering (QA) benchmarks primarily assess short-horizon classification or statistical summaries, largely ignoring the long-term patterns, therapeutic and behavioural contexts, and potential system failures inherent in real-world deployments. To address this, we introduce HealthLoopQA, a comprehensive diagnostic benchmark for evaluating LLM reasoning over continuous diabetes monitoring data. Grounded in a novel taxonomy of eleven atomic reasoning abilities, HealthLoopQA comprises 127 tasks and over 1,500 QA instances spanning process mining, anomaly detection, and prediction over 30-day horizons. To systematically evaluate safety awareness, we complement real-world datasets with a fault-injected simulation testbed modeling diverse device malfunctions and cyber-physical attacks to generate physiologically plausible hazard scenarios. Evaluating state-of-the-art LLMs across prompting and agentic frameworks reveals severe limitations in complex temporal pattern mining. Furthermore, we identify a broader phenomenon of In-context Laziness under long-context prompting, highlighting critical open challenges in deploying LLMs for rigorous long-horizon medical reasoning.

↑