基于检索的开放式评估在哪里失败?从长文本医学答案事实性验证中自动归纳分类体系
Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification
浏览论文内容
中文总结 AI 辅助
本研究通过自动归纳分类体系,揭示基于检索的医学事实性评估在开放式场景中的失败模式,证明模型规模、推理努力和医学微调等改进均无法克服“先检索后验证”范式的根本局限。
中文摘要 AI 辅助
基于检索的事实性评估已成为高风险临床环境中可扩展幻觉检测的主导范式,其中由大语言模型生成的声明需依据权威医学语料库中的证据进行验证。尽管可靠且透明的医学事实验证具有紧迫性,大多数系统仍使用F1等聚合指标来衡量性能,这些指标掩盖了失败发生的位置和原因。现有的RAG诊断方法需要金标准答案或带标注的金标准证据,而这两者在当前场景中均不存在。我们基于开放式MedExpert数据集和3个封闭式数据集的案例研究,引入了两个全面的分类体系,将失败分解为五个质量维度的检索阶段错误,以及六个连续步骤的验证器推理错误。我们采用LLM-as-Judge的自适应模式归纳流程,大规模标注证据质量并分类验证器推理错误,随后在4种检索方法和6个前沿验证器模型上对我们的发现进行压力测试。我们的分析表明,扩大模型规模、增加推理努力、扩展到权威网络来源以及应用医学微调均无法解决这些失败模式,证明它们代表了开放式医学场景中“先检索后验证”范式的根本局限,而非过时系统的产物。我们在该https URL发布代码和数据,以确保结果的完全可复现性。
英文摘要
Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of reliable and transparent medical fact verification, most systems measure performance with aggregate metrics like F1, which obscure where and why failures occur. Existing RAG diagnostics require gold answers or annotated gold evidence, neither of which exists in this regime. We introduce two comprehensive taxonomies, grounded in a case study on the open-ended MedExpert dataset and 3 closed-ended datasets, decomposing failures into retrieval-stage errors along five quality dimensions, and verifier-reasoning errors into six consecutive steps. We adapt an automatic pattern induction pipeline using LLM-as-Judge to label evidence quality and classify verifier reasoning errors at scale, and then stress-test our findings across 4 retrieval methods and 6 frontier verifier models. Our analysis reveals that scaling model size, adding reasoning effort, expanding to authoritative web sources, and applying medical fine-tuning do not resolve these failure modes, demonstrating that they represent fundamental limitations of the retrieve-then-verify paradigm in open-ended medical settings rather than artifacts of outdated systems. We release our code and data at https://anonymous.4open.science/r/Medical_RAG_eval-4AB5 for the full reproducibility of our results.
发表机构
- Johns Hopkins University(约翰霍普金斯大学)
- Center for Language and Speech Processing(语言与语音处理中心)
机构由 AI 辅助整理,请以论文原文为准。