arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CareGraph:一种用于基于证据的个性化纵向健康智能的可审计混合AI框架

CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence

Pratik Ghawate, Tanvi Patil

arXiv 2608.27484首次发表:更新:

AI 中文总结

CareGraph是可审计混合AI框架,用于将异构医疗记录转化为结构化证据,实验显示其在多指标上优于传统方法,为个性化健康系统提供安全受限基础。

AI 中文摘要

人工智能正在变革个性化医疗,但分散的临床、自我报告及可穿戴设备证据仍难以解读和追溯。我们提出CareGraph,一种可审计的混合AI框架,它能将异构记录转化为优先级趋势、缺失上下文指标、受限下一步措施、讨论问题及带有来源链接的解释。CareGraph仅组织证据,不进行诊断、结果预测、治疗选择或自主临床决策。其流程涵盖确定性分析、上下文检测、图构建、受限语言模型合成、证据验证、安全控制及发布门控。测试使用各含400名患者的合成队列进行开发、验证和留存测试。在留存数据上,带有充分性门控的冻结普通最小二乘趋势规则取得0.827准确率、0.837宏F1(95%置信区间为0.819至0.854)及0.974数据不足F1。缺失上下文检测取得0.815严格微F1,而传统检测器仅为0.318。在指定留存基准上,安全规则集1.2版本取得1.000精度、0.950召回率及0.974 F1。一项需对80名患者进行图检索的审计产生79项合成结果和78项展示,无弃权(不执行);1项输出被阻止,1项因无效证据键未通过闭合测试。在56名匹配患者上与整体式GPT-5.6对比,CareGraph速度更快(40.15秒对49.62秒)、内容更短(661词对1163词),且与纵向目标的探索性词汇对齐更好;基线模型使用更少token并引用更多原始证据。图审计验证了来源和确定性检索;图生成的增量效应需配对评估。CareGraph为智能个性化健康系统提供了安全受限的基础。

英文摘要

Artificial intelligence is transforming personalized healthcare, yet fragmented clinical, self reported, and wearable evidence remains difficult to interpret and trace. We present CareGraph, an auditable hybrid AI framework that converts heterogeneous records into prioritized trends, missing context indicators, bounded next steps, discussion questions, and provenance linked explanations. CareGraph organizes evidence without diagnosing, predicting outcomes, selecting treatment, or making autonomous clinical decisions. Its pipeline covers deterministic analysis, context detection, graph construction, constrained language model synthesis, evidence validation, safety controls, and release gating. Tests used synthetic cohorts of 400 patients each for development, validation, and holdout. On holdout data, a frozen ordinary least squares trend rule with a sufficiency gate achieved 0.827 accuracy, 0.837 macro F1 with a 95 percent confidence interval of 0.819 to 0.854, and 0.974 insufficient data F1. Missing context detection achieved 0.815 strict micro F1 versus 0.318 for the legacy detector. On an authored holdout benchmark, safety ruleset version 1.2 achieved 1.000 precision, 0.950 recall, and 0.974 F1. An audit requiring graph retrieval across 80 patients yielded 79 syntheses and 78 presentations without fallback; one output was blocked and one failed closed because of an invalid evidence key. Against monolithic GPT 5.6 on 56 matched patients, CareGraph was faster at 40.15 versus 49.62 seconds, shorter at 661 versus 1,163 words, and showed better exploratory lexical alignment with longitudinal targets; the baseline used fewer tokens and cited more raw evidence. Graph auditing verified provenance and deterministic retrieval; incremental graph effects on generation require paired evaluation. CareGraph offers a safety bounded foundation for intelligent personalized health systems.

Comments21 pages, 7 figures, Code and data: https://github.com/PratikGhawate/ai-personalized-health-intelligence

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑