发表机构
Drexel University; University of Michigan, Ann Arbor(德雷塞尔大学; 密歇根大学安娜堡分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对医学LLM共识形成中一致错误问题,ProbeGuard基于共识形成过程特征和反证探针,实现认证弃权(不执行),显著提升正确与错误共识区分能力并降低选择性风险。
AI 中文摘要
在临床实践中,独立专家之间的共识被视为可靠性的证据,而多轮共识已成为智能体医学问答系统的核心机制。当此类系统需要决定是否信任自身答案时,主要信号再次是共识,即采样答案之间的一致性。然而,共识是正确性的脆弱代理。系统可能一致错误,即每次采样都返回相同的错误答案,而在这些问题上,基于共识的信号不携带任何信息。其原因在于这些信号仅读取共识的最终状态,而丢弃了共识如何形成的过程。通过解决分歧并借助证据达成的共识,与因每个采样共享同一误解而从首个采样起就存在的共识,在最终状态上看起来完全相同。ProbeGuard是一个认证弃权(不执行)框架,其弃权(不执行)决策基于共识的形成方式。过程特征追踪共识轨迹、少数派持续性和检索饱和性。对于一致投票,理由语义熵检查投票背后的理由是否连贯,同时主动探针检索反证并衡量共识是否在反证下存续。随后,分层“先学习后测试”校准将这些分数转换为选择性风险的无分布界限。我们在三个医学问答基准和一个困难前沿参考上,基于已发布的多轮智能体RAG基底,针对六个弃权(不执行)基线评估了ProbeGuard。在MedQA上,13.4%的一致投票是错误的,且没有任何基于共识的信号能标记它们。过程信号将正确与错误共识的区分能力从随机水平提升至0.696 AUROC。认证规则在一致层问题中回答了十分之六的问题,观测选择性风险为9.0%,而一旦域内校准数据积累,则能回答十分之九的问题。
英文摘要
In clinical practice, agreement among independent experts is treated as evidence of reliability, and multi-round consensus has become a core mechanism of agentic medical question-answering systems. When such a system must decide whether to trust its own answer, the prevailing signal is again agreement, now among the sampled answers. But agreement is a fragile proxy for correctness. A system can be unanimously wrong, returning the same incorrect answer on every sample, and on these questions agreement-based signals carry no information. The cause is that these signals read only the final state of the consensus and discard how it was reached. Agreement that was reached by resolving disagreement with evidence looks identical, at the end, to agreement that was present from the first sample because every sample shares one misconception. ProbeGuard is a certified abstention framework that bases the abstention decision on how the consensus formed. Process features trace agreement trajectories, minority persistence, and retrieval saturation. For unanimous votes, rationale semantic entropy checks whether the reasons behind the vote cohere, and an active probe retrieves counter-evidence and measures whether the consensus survives. A stratified Learn-then-Test calibration then converts these scores into a distribution-free bound on selective risk. We evaluate ProbeGuard on three medical QA benchmarks and a hard-frontier reference, with a published multi-round agentic RAG substrate, against six abstention baselines. On MedQA, 13.4% of unanimous votes are wrong, and no agreement-based signal can flag them. Process signals raise the discrimination of correct from incorrect consensus from chance to 0.696 AUROC. The certified rule answers six in ten unanimous-layer questions at an observed selective risk of 9.0%, and nine in ten once in-domain calibration data accumulate.
CommentsAccepted at the GenAI4Health Workshop at NeurIPS 2026