验证、修复、重复还是停止?大语言模型代理中用于有噪声验证-修复循环的稳健停止方法
Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents
浏览论文内容
中文总结 AI 辅助
研究大语言模型代理中验证-修复循环有噪声时缺乏停止原则的问题,提出VRR-Stop框架,通过四参数噪声模型和信念过滤估计有效性,依真实边际增益符号决策,配对VRR-Guard,提升了停止可靠性与真实有效性。
中文摘要 AI 辅助
验证-修复循环是大语言模型代理在代码生成、数学推理和工具使用中纠正错误计划的标准手段。当验证器和修复器都有噪声时,修复可能会破坏已正确的计划,报告的接受率持续上升而实际有效性下降,现有方法缺乏决定修复何时停止的原则基础。我们提出VRR-Stop,一种用于有噪声验证-修复-重复(VRR)循环的稳健停止框架。一个四参数噪声模型区分验证器的错误接受和错误拒绝与修复器的修复和破坏行为。信念过滤将重复的验证投票转化为对确定有效性的估计,循环根据真实边际增益的符号进行提交或修复,这仅需要符号可识别性而非所有参数的精确恢复。当验证器的辨别力接近零时,校准本身会失败且估计误差可能会翻转停止符号,所以我们将VRR-Stop与VRR-Guard配对,VRR-Guard是一种无估计的后备方法,仅在有足够验证余量时才替换现有候选方案。在GSM8K压力设置下,VRR-Stop比固定的五轮修复将最终真实有效性提高了60.6个百分点,平均成本为0.72轮修复。在各种设置中,停止可靠性由验证器辨别力和决策余量共同决定,而非估计误差的绝对大小。
英文摘要
Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool use. When both the verifier and the repairer are noisy, repair can damage already-correct plans, and reported acceptance keeps rising while true validity falls, so existing methods lack a principled basis for deciding when repair should stop. We propose VRR-Stop, a robust stopping framework for noisy verify-repair-repeat (VRR) loops. A four-parameter noise model separates verifier false acceptance and false rejection from the repair and damage behavior of the repairer. Belief filtering turns repeated verification votes into an estimate of committed validity, and the loop commits or repairs according to the sign of the true marginal gain, which requires only sign identifiability rather than accurate recovery of all parameters. When verifier discrimination approaches zero, calibration itself fails and estimation error can flip the stopping sign, so we pair VRR-Stop with VRR-Guard, an estimation-free fallback that replaces the incumbent candidate only under a sufficient verification margin. On a GSM8K stress setting, VRR-Stop improves final true validity by 60.6 percentage points over fixed five-round repair at an average cost of 0.72 repair rounds. Across settings, stopping reliability is governed jointly by verifier discrimination and the decision margin rather than by the absolute size of estimation error.