MedSNIP:面向医学事实核查的片段级粒度构建与基准测试
MedSNIP: Building and Benchmarking Snippet-Level Granularity for Medical Fact Verification
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
- Yale University(耶鲁大学)
- Korea University(高丽大学)
- Beijing Normal University(北京师范大学)
- INSAIT, Sofia University, “St. Kliment Ohridski”(索非亚大学INSAIT研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对原子级分解破坏临床结构的问题,提出MedSNIP片段级验证方法及基准,通过保留局部结构提升假类F1并减少核查调用。
AI中文摘要:
医学声明的正确性往往不仅取决于声明本身,还取决于其周围的临床结构。一个声明可能需要实验室参考范围、因果或条件关联,或患者特定细节才能被正确判断,而原子级分解可能会割裂这些依赖关系,使核查器面对临床信息不完整的声明。我们将医学事实核查重新表述为片段级验证,其中从句分组单元保留了局部临床结构。我们引入了MedSNIP-Bench,一个用于片段级医学事实核查的人工标注基准,以及MedSNIP,一个自动片段生成流水线。MedSNIP-Bench涵盖276个消费者健康和临床小插曲响应,被分割为2,524个片段,带有双重通用和患者情境标签以及六种结构模式代码。MedSNIP在MedSNIP-Bench上针对人类片段边界进行评估,然后用于为外部语料库生成片段级单元。在MedSNIP-Bench、HealthFC和MedHallu上,片段级验证保持或提高了假类F1分数,其增益集中在答案足够长以至于可被分割以及核查器足够强大以利用恢复结构的情况下。最大的合并模式增益出现在因果-条件临床链上。它还将核查器调用减少了24-73%,尽管这种节省仅在分解成本低廉时才能端到端保留,而开放权重分解器使得在无分块保真度损失的情况下实现这一点成为可能。
英文摘要:
A medical claim's correctness often depends not on the claim alone, but on the clinical structure around it. A claim may require a lab reference range, a causal or conditional link, or patient-specific details to be judged correctly, and atom-level decomposition can fragment these dependencies, leaving the verifier with clinically incomplete claims. We reformulate medical fact-checking around snippet-level verification, where clause-grouped units preserve local clinical structure. We introduce MedSNIP-Bench, a human-annotated benchmark for snippet-level medical fact verification, and MedSNIP, an automatic snippet-generation pipeline. MedSNIP-Bench covers 276 consumer-health and clinical-vignette responses, segmented into 2,524 snippets with dual in-general and in-patient-context labels and six structural pattern codes. MedSNIP is evaluated against human snippet boundaries on MedSNIP-Bench and then used to generate snippet-level units for external corpora. Across MedSNIP-Bench, HealthFC, and MedHallu, snippet-level verification preserves or improves false-class F1, with gains concentrated where answers are long enough to fragment and where the verifier is strong enough to exploit the recovered structure. The largest merge-pattern gain is on causal-conditional clinical chains. It also reduces verifier calls by 24-73%, though the saving survives end-to-end only when decomposition is cheap, which an open-weight decomposer makes possible at no loss of chunking fidelity.