发表机构
Tynapse(泰纳普斯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出评审价值模型,结合错误程度、干预可行性、影响和成本来优化有限预算下的答案评审排序,通过WAER和PRRE指标在TAT-QA/SciFact基准上验证,显著降低修复后残余暴露。
AI 中文摘要
大语言模型助手通常会产生比人类在用户看到答案之前所能评审的更多的答案。大多数评估关注答案是否错误、缺乏依据或置信度低。而有限评审预算则提出不同的问题:在固定的评审预算下,应优先检查哪些答案。仅凭风险是不够的:高风险答案可能难以修复,而中等风险答案可能直接从现有证据中得到纠正。对于生成答案的评估,我们将评审优先级建模为暴露减少,其中评审价值结合了估计的错误程度、干预可行性、影响和成本。我们使用错误答案暴露率(WAER)——未被评审的错误答案比例,以及修复后残余暴露(PRRE)——在确定性基准支持的修复后仍暴露的比例,来评估评审队列。PRRE使用可修复性规则,这些规则在数值上不重复用于排序的可行性分数。在包含720个条目的TAT-QA/SciFact压力测试基准上,评审价值排序在20%预算下使答案级WAER几乎保持不变(0.605对0.600),但将PRRE从0.881降至0.716。这些结果表明,可信赖的大语言模型评估不仅应衡量错误检测,还应衡量有限的评审能力如何减少暴露的错误答案。
英文摘要
LLM assistants often produce more answers than humans can review before users see them. Most evaluations ask whether an answer is wrong, unsupported, or low-confidence. Bounded review budgets instead ask which answers should be checked first under a fixed review budget. Risk alone is not enough: a high-risk answer may be hard to repair, while a moderately risky answer may be directly correctable from available evidence. For generated-answer evaluation, we model review prioritization as exposure reduction, where review value combines estimated wrongness, intervention affordance, impact, and cost. We evaluate review queues with Wrong-Answer Exposure Ratio (WAER), the fraction of wrong answers left unreviewed, and post-repair residual exposure (PRRE), the fraction still exposed after deterministic benchmark-supported repairs. PRRE uses repairability rules that do not numerically reuse the affordance scores used for ranking. On a 720-item TAT-QA/SciFact stress benchmark, review-value ranking keeps answer-level WAER nearly unchanged at 20% budget (0.605 vs. 0.600) but lowers PRRE from 0.881 to 0.716. These results show that trustworthy LLM evaluation should measure not only error detection, but also how limited review capacity reduces exposed wrong answers.
CommentsAccepted at SeT-LLM 2026 Workshop, KDD 2026