arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13475cs.AI

GAVEL:跨提取时间线差异的基于案例报告的裁决——比较临床时间线与病例报告

Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports

发表机构普林斯顿大学 · 国家医学图书馆 · 圣路易斯华盛顿大学
查看机构详情
  • Princeton University(普林斯顿大学)
  • National Library of Medicine(国家医学图书馆)
  • Washington University in St. Louis(圣路易斯华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Jack Cummins, Sayantan Kumar, Ketan Tamirisa, Jeremy C. Weiss

首次发表
浏览论文内容

中文总结 AI 辅助

GAVEL是一种LLM法官协议,通过比较时间线与病例报告来裁决差异,无需金标准,经评估将差异从每份报告7.63降至0.85,并在77.0%比较中优先选择合并时间线。

中文摘要 AI 辅助

现有的从病例报告中提取临床时间线的流程使用专家参考进行评估,但受到参考标注不完善和事件对齐不精确的限制。我们开发了GAVEL,一种LLM法官协议,它将两条时间线与病例报告进行比较,并为每个差异返回差异类型、裁决结果和报告段落。我们评估了事件匹配器,审查了来自GPT5.6sol和DeepSeek V3.2的2,738项发现,对六个LLM提取器和两名人工标注者进行了排名,并测试了GAVEL引导的合并。真实匹配率在0.10阈值正下方为60%,正上方为48%。人工审查确认了89.4%和88.6%的发现。在126份报告中,合并后的时间线在77.0%的比较中被优先选择(95%置信区间,69.8%至84.1%),并将归因于被评估时间线的差异从每份报告7.63降至0.85。GAVEL支持基于报告的比较和修订,而无需将任一时间线视为金标准。

英文摘要

Existing pipelines for clinical timeline extraction from case reports are evaluated using an expert reference and are limited by imperfect reference annotations and imprecise event alignment. We developed GAVEL, an LLM judge protocol that compares two timelines with the case report and returns a discrepancy type, verdict, and report passage for each difference. We evaluated the event matcher, reviewed 2,738 findings from GPT5.6sol and DeepSeek V3.2, ranked six LLM extractors and two human annotators, and tested GAVEL guided merging. True match rates were 60% immediately below and 48% immediately above the 0.10 cutoff. Manual review confirmed 89.4% and 88.6% of findings. Across 126 reports, merged timelines were preferred in 77.0% of comparisons (95% CI, 69.8 to 84.1%) and reduced discrepancies attributed to the evaluated timeline from 7.63 to 0.85 per report. GAVEL supports report-based comparison and revision without treating either timeline as ground truth.

补充信息

↑