arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

验证陷阱:理解代码生成中错误前提下的测试时选择失败

Verification Trap: Understanding Test-Time Selection Failures under False Premises in Code Generation

Feng He, Hejia Wang, Linghao Meng, Ming Gao, Qiankun Li

arXiv 2610.05170首次发表:更新:

发表机构

University of Science and Technology of China; Beijing University of Posts and Telecommunications; National University of Singapore; Nanyang Technological University(中国科学技术大学; 北京邮电大学; 新加坡国立大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示代码生成中生成器与验证器共享错误前提导致测试时选择失败的“验证陷阱”,通过三个基准和五个模型验证其影响,并提出基于验证器可见特征的无金标预测器(AUROC 0.846)及解耦证据作为缓解方向。

AI 中文摘要

测试时计算已成为改进代码生成的核心方式:系统采样多个候选程序,并利用验证器可见的证据来选择最终输出。这一范式隐含地假设验证器提供了独立于生成器的纠正信号。我们在误导性任务前提下挑战这一假设。当生成器和验证器共享一个错误前提时,它们会通过一个错误的信念而耦合:生成器产生与前提一致的捷径,而验证器提供的证据却无法暴露这些捷径。因此,即使池中存在隐藏测试正确的程序,选择器也可能选择隐藏测试错误的候选程序。我们将这种失败模式称为“验证陷阱”。在三个代码生成基准和五个代码模型上,错误前提持续降低了首样本正确性,在64样本测试时选择后降低了选择器选择的正确性,并放大了可恢复的错误选择。从机制上讲,验证器编写的测试继承了前提层面的盲点,将验证器可见的候选空间重塑为偏离隐藏测试正确性。这些痕迹使得验证陷阱在隐藏执行之前可预测:一个使用验证器可见特征的轻量级无金标预测器达到了0.846的AUROC。我们的结果将解耦证据确定为一个关键的缓解轴:耦合扩展提供的恢复有限,而前提无关的鲁棒性审计器恢复了大量的神谕上限。

英文摘要

Test-time compute has become a central way to improve code generation: systems sample multiple candidate programs and use verifier-visible evidence to select the final output. This paradigm implicitly assumes that the verifier provides a corrective signal independent from the generator. We challenge this assumption under misleading task premises. When the generator and verifier share a false premise, they become coupled through a mistaken belief: the generator produces premise-consistent shortcuts, while the verifier supplies evidence that fails to expose them. Consequently, the selector may choose a hidden-test-wrong candidate even when a hidden-test-correct program exists in the pool. We call this failure mode Verification Trap. Across three code-generation benchmarks and five code models, false premises consistently degrade first-sample correctness, reduce selector-chosen correctness after 64-sample test-time selection, and amplify recoverable mis-selection. Mechanistically, verifier-written tests inherit the premise-level blind spot, reshaping verifier-visible candidate space away from hidden-test correctness. These traces make Verification Trap predictable before hidden execution: a lightweight gold-free predictor using verifier-visible features reaches 0.846 AUROC. Our results identify decoupled evidence as a key mitigation axis: coupled scaling provides limited recovery, whereas premise-agnostic robustness auditors recover substantial oracle headroom.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑