重新思考大语言模型验证:证据结构、不确定性与选择性优化
Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement
浏览论文内容
中文总结 AI 辅助
该研究针对LLMs医疗应用的安全问题,提出两阶段框架,利用模型弃权信号优化推理,在GPT-5.5、DeepSeek-R1模型及MedReason、MedQA数据集上显著提升了医疗假设验证的准确率。
中文摘要 AI 辅助
大语言模型(LLMs)常依赖捷径而非系统推理,引发医疗应用中的安全担忧。允许模型在不确定时弃权(不执行)可提升可靠性,但会引入覆盖率与准确率的权衡。我们针对多项选择场景下的医疗假设验证,提出两阶段框架,仅在模型弃权时应用靶向本体接地以管理该权衡。研究表明,弃权并非随机,而是反映真实不确定性,弃权预测与较低置信度相关。在两个前沿模型(通过Azure OpenAI API访问的GPT-5.5与DeepSeek-R1)上,该框架使问题级准确率提升9.6个百分点(从82.9%升至92.5%),假设级准确率提升4.2个百分点(从92.0%升至96.2%)。在MedReason与MedQA上开展的实验显示,弃权可被重新用作选择性推理优化的控制信号,无需显式构建知识图谱即可达到知识图谱级性能。
英文摘要
Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models to abstain when uncertain improves reliability but introduces a coverage accuracy tradeoff. We propose a two-stage framework for medical hypothesis verification in multiple-choice settings that manages this tradeoff through targeted ontology grounding, applied only when the model abstains. We show that abstention is not random but reflects genuine uncertainty, with abstained predictions associated with lower confidence. Across two frontier models (GPT-5.5, accessed via the Azure OpenAI API, and DeepSeek-R1), the proposed framework improves question-level accuracy by 9.6 percentage points (82.9% to 92.5%) and hypothesis-level accuracy by 4.2 percentage points (92.0% to 96.2%). Our experiments conducted on MedReason and MedQA show that abstention can be repurposed as a control signal for selective reasoning refinement, achieving knowledge-graph-level performance without explicit knowledge graph construction.
发表机构
- Indian Institute of Technology Jammu(贾姆穆印度理工学院)
- Microsoft Research(微软研究院)
- SCB Dental College and Hospital(SCB牙科学院与医院)
机构由 AI 辅助整理,请以论文原文为准。