arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04611cs.SEcs.AI

顺序是保障:基于静态优先学习提议的验证者预算代码删除

The Order Is the Guarantee: Verifier-Budgeted Code Deletion with Static-First Learned Proposals

  • University of Hong Kong(香港大学)
  • Zhejiang University(浙江大学)
  • Dalian University of Technology(大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Han Wang, Jie Li, Ru Zhang

AI总结:

该研究针对验证者预算有限时的AI代码删除问题,提出DELSCOUT调度方法,结合静态与学习候选排序,提升已验证删除覆盖率并控制验证器调用,明确了模型、顺序与执行的分工。

AI中文摘要:

前沿编码模型现在在编程基准上达到或超过了优秀人类参考水平,但基准测试成功并不意味着软件具备可维护性。提示驱动的“氛围编码”是增量式的:新分支、防护和回退的积累速度快于过时逻辑的移除速度。我们研究逆问题:当执行验证能力有限时,AI系统应如何删除代码。我们将冗余代码缩减表述为提议调度:排序器对单语句删除候选进行排序,执行套件接受首个通过的候选,且预算限制了可测试的候选数量。我们的核心观察是,候选顺序而非模型置信度,是部署系统可调控的控制面。DELSCOUT实现了两种调度方案:给定代表性目标域验证,5个插槽的预算将3个插槽分配给确定性最短优先候选,2个插槽分配给互补学习候选;在9次MBPP复现中,使用0.5B、0.6B和8B规模的排序器,这使已验证删除覆盖率相对提升9.5%(增加6.7个已接受任务),同时消耗的验证器调用略少于匹配的静态基线。若没有此类验证,相同排序器在分布偏移下可能会损失覆盖率,因此我们改为先评估完整静态前缀,仅在之后附加学习候选;对于确定性验证器,这使得覆盖率和代码量缩减在构造上呈非递减趋势,验证器调用的测量增幅为4.8%至62.5%。MBPP+则消除了域内优势,表明调度控制搜索过程,而仅测试套件决定“保留行为”的含义。最终形成可审计的分工:模型拓展可删除代码的搜索范围,顺序限制排序错误提议可能造成的损害,执行环节对每一项已提交的删除保留控制权。

英文摘要:

Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply maintainable software. Prompt-driven "vibe coding" is additive: new branches, guards, and fallbacks accumulate faster than obsolete logic is removed. We study the inverse problem-how an Al system should remove code when execution-verification capacity is finite. We formulate redundant-code reduction as proposal scheduling: a ranker orders single-statement deletion candidates, an execution suite accepts the first candidate that passes, and a budget bounds how many candidates may be tested. Our central observation is that candidate order, not model confidence, is the control surface a deployment can reason about. DELSCOUT instantiates two schedules. Given representative target-domain validation, a five-slot budget spends three slots on deterministic shortest-first candidates and two on complementary learned candidates; across nine MBPP replications with 0.5B, 0.6B, and 8B rankers this raises verified-deletion coverage by 9.5% relative (+6.7 accepted tasks) while consuming slightly fewer verifier calls than the matched static baseline. Without such validation the same rankers can lose coverage under shift, so we instead evaluate the complete static prefix first and append learned candidates only afterwards; for a deterministic verifier this makes coverage and character reduction non-decreasing by construction, at a measured 4.8-62.5% increase in verifier calls. MBPP+ then erases the in-domain advantage, showing that scheduling governs search while the test suite alone governs what "preserving behavior" means. The result is an auditable division of labor: models widen the search for removable code, order bounds the damage a mis-ranked proposal can do, and execution retains authority over every committed deletion.

↑