arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FairGlucose:揭示人群水平验证中隐藏亚组差异的CGM公平性基准

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao

arXiv 2608.18296首次发表:更新:

发表机构

Carey Business School; Johns Hopkins University; Welldoc, Inc.(凯瑞商学院; 约翰斯·霍普金斯大学; 韦尔多克公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究构建FairGlucose基准,发现CGM预测模型人群水平验证掩盖亚组差异,T1D患者误差更高,前沿LLM表现逊于专业神经模型,推动数字健康AI采用亚组分层公平评估标准。

AI 中文摘要

随着基于CGM(连续血糖监测)的AI工具接近临床部署,其准确性是否在不同患者人口统计群体间保持公平性尚未得到充分验证。为实现该评估,我们构建了FairGlucose,这是一个包含300名患者的CGM队列,在12个人口统计层(年龄×性别×1型/2型糖尿病)间保持平衡,其中包含132480个预测样本,以及81名患者记录的3945个独特行为事件(饮食、运动、药物)。我们对四个模型家族的33个模型进行2小时血糖预测的基准测试,发现人群水平的外部验证会掩盖大量亚组差异。总体分布外指标看似稳定(约1.0),但亚组水平的比率范围为0.8至1.4,1型糖尿病(T1D)患者的预测误差比2型糖尿病(T2D)患者高6mg/dL(p<0.001)。这种差异在所有33个模型中均存在,表明是预测任务的属性而非单一架构的问题。进一步分析显示,亚组性能差距与临床困难病例的比例一致,且不同人口统计群体的输入长度敏感性存在差异,这推动了个性化配置。前沿大语言模型(LLM)的表现比专业神经模型差1-6mg/dL;即使在获得事件的理想访问权限下,行为事件的贡献也可忽略不计(约0.1mg/dL)。这些发现表明,仅人群水平验证不足以评估数字健康AI的公平性,推动将亚组分层报告作为默认标准。

英文摘要

As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows that subgroup performance gaps align with the proportion of clinically hard cases, and that input-length sensitivity varies across demographics, motivating personalized configurations. Frontier LLMs underperform specialized neural models by 1-6 mg/dL; behavioral events contribute negligibly (approximately 0.1 mg/dL) even under oracle event access. These findings establish that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑