AI 中文总结
该研究针对LLM代码审查,测试4种审查界面的策略与概率分离情况,提出模块化流水线可提升概率准确性、降低损失,验证了风险与行动分离评估的必要性。
AI 中文摘要
大语言模型(LLM)代码审查工具通常在一个提示中同时估算补丁风险并做出批准决策。概率应基于证据,而成本应决定基于该证据采取的行动。我们测试了4种已部署的审查界面是否保持了这种分离,使用了针对720个候选补丁的15792条响应,每个补丁对应360个仓库问题,其中每个问题都有一个通过存档测试工具的补丁和一个未通过的补丁。在补丁和监控证据固定的匹配调用中,将等成本策略替换为10:1的误接受策略,会使报告的失败概率平均变化13.6至16.9个百分点。对于每个审查工具,高成本提示下返回的行动比拒绝所有补丁更差。将相同的高成本规则应用于等成本下得出的概率,可降低所有4个系统的损失,表明概率引出本身会导致额外损失。我们还评估了一种模块化流水线,该流水线在无策略信息的情况下引出风险,结合独立的监控评分,并在代码中应用成本。与校准后的仅审查工具评分相比,该流水线提高了平均概率准确性,且在等成本下,每个问题的平均损失降低0.073,同时接受58%至68%的补丁;在10:1的比例下,它不接受任何补丁,与全部拒绝一致。因此,下游策略可能会改变其本应使用的概率,这促使对风险、外部证据和行动进行单独评估。
英文摘要
LLM code reviewers often estimate patch risk and make approval decisions in one prompt. A probability should depend on evidence; costs should determine the action taken from it. We test whether four deployed reviewer interfaces preserve this separation using 15,792 responses on 720 candidate patches, with one that passed and one that failed an archived test harness for each of 360 repository issues. In matched calls with the patch and monitor evidence fixed, replacing an equal-cost policy with a 10:1 false-accept policy changes reported failure probabilities by 13.6 to 16.9 percentage points on average. For every reviewer, the actions returned under the high-cost prompt are worse than rejecting all patches. Applying the same high-cost rule to probabilities elicited under equal costs reduces loss for all four systems, showing that probability elicitation itself contributes to the excess loss. We also evaluate a modular pipeline that elicits risk without policy information, combines an independent monitor score, and applies costs in code. Relative to calibrated reviewer-only scores, the pipeline improves average probability accuracy and, at equal costs, reduces mean loss by .073 per issue while accepting 58 to 68% of patches. At 10:1, it accepts none and matches reject-all. Downstream policy can therefore change the probability it is meant to use, motivating separate evaluation of risk, outside evidence, and action.
Comments20 pages, 6 figures; includes technical appendices. Code and data: https://github.com/rasvik/when-policies-change-probabilities