发表机构
University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究公开测试通过后代码隐藏错误监控问题,引入代码监控红队协议并实例化为代码监控基准。通过大量实验发现弱验证器虽能改进但仍易漏错,对抗性压力影响验证效果,GLM - 5.1验证器缩小部分差距,揭示了验证器故障与证据限制的问题。
AI 中文摘要
可见测试是大语言模型生成代码的常见关卡,但通过这些测试并不能保证规范的正确性。我们研究了一个类似部署的监控问题:在代码通过公开测试后,一个较弱的大语言模型验证器能否识别出残留的隐藏错误?我们引入了代码监控红队,这是一种监控红队协议,它在改变生成器压力、验证器框架和从弱到强的能力时,固定公共检查信息边界。我们将其实例化为代码监控基准,涵盖函数级、数据科学和工作流代码。在71000个生成的候选代码中,43677个通过了公开测试,其中23081个未能通过隐藏测试。弱验证器通过框架和模型家族得到改进,但在5%的误报率下仍会遗漏大多数隐藏错误。作为一种鲁棒性压力测试,对抗性的公开测试过拟合压力会降低验证器的曲线下面积,并在大多数单元中提高低误报率下的遗漏率。一个GLM - 5.1验证器在相同证据边界下缩小了部分差距;可推断性审计表明,剩余的遗漏将验证器故障与M1证据限制混在了一起。
英文摘要
Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-like monitoring problem: after code has passed public tests, can a weaker LLM verifier identify the residual hidden bugs? We introduce Code Monitor Red Teaming, a monitor-red-teaming protocol that fixes a public-check information boundary while varying generator pressure, verifier scaffolding, and weak-to-strong capability. We instantiate it as CodeMonitorBench, spanning function-level, data-science, and workflow code. Across 71,000 generated candidates, 43,677 pass public tests and 23,081 of those fail hidden tests. Weak verifiers improve with scaffolding and model family, but still miss most hidden bugs at 5% false-positive rate. As a robustness stress test, adversarial public-test-overfit pressure lowers verifier AUROC and raises low-FPR miss rates in most cells. A GLM-5.1 verifier recovers part of the gap under the same evidence boundary; an inferability audit shows that remaining misses mix verifier failures with M1 evidence limits.