arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JuryProbe:用于将无参考事实性评判小组路由至基于事实的验证的经验共识风险诊断方法

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

Tianxin Zhou, Ruixi Lin

arXiv 2608.20607首次发表:更新:

AI 中文总结

针对无参考事实性LLM评判小组的共识风险诊断方法JuryProbe,通过校准路由策略减少假阴性盲区导致的错误,在FEVER数据上验证了其有效性,可降低假接受并减少参考获取。

AI 中文摘要

由低成本大语言模型(LLM)评判员组成的小组越来越多地做出接受或升级的决策。在事实性评估场景中,因多个无参考评判员达成一致而接受某一主张,可能会产生隐藏风险:这种一致可能反映出共同的假阴性盲区,而非独立证据。我们提出JuryProbe,一种针对无参考事实性评判小组的经验共识风险诊断方法,搭配基于校准的路由策略。JuryProbe利用仅含假阴性(FN-only)的评判员相关性和假共识提升值,从带标签的校准探针中估计共识风险;当标记为高风险时,将无参考多数接受结果路由至拥有可信参考的相同评判员。在经审计的FEVER数据损坏情况下,无参考小组表现出相关的假阴性(FN-only相关性为0.402和0.368;提升倍数为3.13倍和18.13倍),而在可信参考最佳情况诊断下,针对最小对和非最小对证据的一致假共识均降至零。在标记场景中,该路由策略按构造等价于对每一项无参考多数接受结果进行基于事实的验证(在34/34的拆分中得到验证);改进来自接受条件下的基于事实的验证,而诊断则决定是否激活该机制。一个固定的预先指定规则在合成、基准生成和科学类别中,每10个拆分中有8-10个被标记,在阴性对照中每10个拆分中有0个被标记,此时弃权(不执行)可减少28%的参考获取,同时假接受仅增加0.004。在弱BM25检索下,假接受减少仍持续存在,但会带来显著的覆盖成本,而过时的弃权标签需要定期重新校准。JuryProbe不提供正式的风险保证,也未在自然小组上建立可靠的弃权机制;其支持的贡献是对高风险小组错误依赖的经验诊断。

英文摘要

Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind spots rather than independent evidence. We introduce JuryProbe, an empirical consensus-risk diagnostic for reference-free factuality judge panels, paired with a calibration-based routing policy. JuryProbe estimates consensus risk from a labeled calibration probe using false-negative-only (FN-only) judge correlation and false-consensus lift; when flagged high-risk, reference-free majority accepts are routed to the same judges with trusted references. On audited FEVER corruptions, reference-free panels show correlated false negatives (FN-only correlations 0.402 and 0.368; lifts 3.13x and 18.13x), while unanimous false consensus drops to zero under a trusted-reference best-case diagnostic on both minimal-pair and non-minimal-pair evidence. In flagged settings, the routed policy is by construction equivalent to grounding every reference-free majority accept (verified in 34/34 splits): improvement comes from accept-conditioned grounding, while the diagnostic determines whether to activate it. A fixed, pre-specified rule flags 8-10 of 10 splits across synthetic, benchmark-authored, and scientific families and 0 of 10 on a negative control, where standing down avoids 28% of reference acquisitions at a 0.004 increase in false accepts. False-accept reduction persists under weak BM25 retrieval at substantial coverage cost, while stale stand-down labels require periodic recalibration. JuryProbe provides no formal risk guarantee and does not establish reliable stand-down on natural panels; its supported contribution is an empirical diagnostic of high-risk panel error dependence.

CommentsAccepted at Transactions on Machine Learning Research (TMLR), 08/2026. 21 pages, 1 figure, 16 tables

Journal refTransactions on Machine Learning Research, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑