先验审计-修复上下文调整大语言模型验证器阈值以趋向宽松
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
浏览论文内容
中文总结 AI 辅助
该研究发现,LLM检查器上下文中的审计→修复事件会降低误报率,调整了验证器阈值趋向宽松,且该效应与负性不对称性预测相反,不随推理启用而消失。
中文摘要 AI 辅助
自动化检查流程越来越多地将一个语言模型作为检查器,另一个语言模型(或同一个)作为修复器。我们探究这种架构是否会改变检查器的报告结果。在对人工验证正确的ProcessBench轨迹进行测量时,保持当前任务的字节级完全相同,我们发现模型上下文中已完成的审计→修复事件,在15种模型×措辞组合中都降低了误报率,与长度匹配的非审计对照组相比,降幅为2.8至11.5个百分点,相对于该对照组的减少幅度为9%至25%。这一方向与累积消息文献的预测相反:在审计报告错误的事件中,在该操作顺利的模型的全部5种措辞上,误报率会进一步降低,尽管负性不对称性预测会有更多标记。对该事件的分解发现,修复内容和审计判断具有互补性:不同组件对不同模型家族产生影响。信号检测分析将变化定位在阈值而非辨别力上——15种组合中有15种的标准发生了变化,且在13种组合中经校正后仍保持有效,而d'则未保持,尽管d'测试的灵敏度天生只有一半;对50个误报的人工审计发现,82%的误报完全错误,因此在该操作点上,这种调整不一定有害。启用推理后,该效应在两个测试模型上都保持了相对规模,阈值读数也依然成立。
英文摘要
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.
发表机构
- University of California, Santa Cruz(加州大学圣克鲁兹分校)
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。