arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

信任前审核:面向AI辅助个人健康记录的基于来源的完整性闸门

Review Before Trust: Source-Grounded Integrity Gates for AI-Assisted Personal Health Records

Nora Girda, Adrian Groza

arXiv 2608.29965首次发表:更新:

发表机构

Technical University of Cluj-Napoca(克卢日-纳波卡技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对AI辅助个人健康记录的完整性风险,提出Medical DataCloud中的证据门控信任提升模型,通过源文档验证机制控制生成数据的使用,经测试验证了技术可行性。

AI 中文摘要

大型语言模型可将医疗文档转换为结构化数据,但看似合理的输出可能缺乏来源支持。将此类输出保留在随时间累积患者信息的纵向健康记录中,会产生完整性风险:未经验证的数据可能影响后续的汇总、趋势分析或预防性护理计算。我们提出一种证据门控的信任提升模型,该模型会将生成的数据保持为临时状态,直至确定性监控器对照源文档对其进行验证。仅当源文档包含唯一支持性引述、相关字段位于同一实验室行内且保留了所需来源时,监控器才会批准候选数据用于指定下游用途。生成器无法批准自身输出,缺失或模糊的证据会导致拒绝,被拒绝的候选数据会保留供人工审核,而非被悄悄丢弃。我们在个人健康记录应用Medical DataCloud中实现了该模型,并通过自动化测试和保存的提取输出重放对其进行评估。所有22项一致性和变异测试均通过。重放涵盖9份包含102个人工标注行的历史实验室PDF报告,这些报告生成97个数值候选数据:架构验证接受全部97个,早期的数据包级证据检查接受94个,强化后的引述与行级策略批准72个,同时保留25个供审核。本研究评估的是系统完整性,而非临床正确性或临床安全性。结果表明,可执行边界在防止生成的声明授权其自身在纵向健康记录中重复使用方面具有技术可行性。

英文摘要

Large language models can convert medical documents into structured data, but plausible output may still be unsupported by the source. Persisting such output in a longitudinal health record, a record that accumulates patient information over time, therefore creates an integrity risk: unverified data may influence later summaries, trends, or preventive-care computations. We introduce an evidence-gated trust-promotion model that keeps generated data provisional until a deterministic monitor verifies it against the source document. The monitor admits a candidate for a specified downstream use only when the source contains a unique supporting quotation, the relevant fields occur within the same laboratory row, and the required provenance is preserved. The generator cannot approve its own output, missing or ambiguous evidence causes refusal, and refused candidates remain available for human review rather than being silently discarded. We implement the model in Medical DataCloud, a personal health-record application, and evaluate it through automated tests and a replay of saved extraction outputs. All 22 conformance and mutation tests pass. The replay covers nine historical laboratory PDF reports containing 102 manually labelled rows. The reports produce 97 numeric candidates: schema validation accepts all 97, an earlier packet-level evidence check accepts 94, and the hardened quotation- and row-level policy admits 72 while retaining 25 for review. The study evaluates system integrity rather than clinical correctness or clinical safety. The results demonstrate the technical feasibility of an enforceable boundary that prevents generated claims from authorizing their own reuse in a longitudinal health record.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑