arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SCATE:学习监督编码代理以实现经济高效的测试生成

SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation

Sijia Gu, Noor Nashid, Ali Mesbah

arXiv 2607.08983首次发表:更新:

发表机构

University of British Columbia(不列颠哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对编码代理惰性生成致代码覆盖不足问题,提出SCATE框架,将监督设为上下文博弈问题,依覆盖和可测试性指标选测试动作,集成不同编码代理,相比基线和其他方法有更高覆盖,能动态优化代理优势。

AI 中文摘要

虽然自主编码代理在自动化测试生成方面取得了显著进展,但它们仍受惰性生成的根本限制,即代理过早终止任务并系统地避开复杂程序逻辑,导致代码覆盖不足。目前缓解这种过早终止需要持续的人工监督,这形成了瓶颈。我们提出SCATE,一个用于对编码代理进行自适应、自动监督的框架,在测试生成期间取代人工干预。通过将监督表述为上下文博弈问题,SCATE基于当前覆盖和类可测试性指标学习选择最有前景的测试动作,在最小化浪费的生成努力的同时最大化覆盖增益。实证评估表明,SCATE与不同编码代理无缝集成。应用于GEMINI-CLI时,它比仅代理基线实现了32.3%更高的行覆盖和30.9%更高的分支覆盖。与CLAUDE CODE的比较证实该框架动态调整其策略以优化每个代理的独特优势。SCATE在所有指标上也始终优于最先进的非代理方法。

英文摘要

While autonomous coding agents have significantly advanced automated test generation, they remain fundamentally limited by lazy generation, a phenomenon where agents prematurely terminate tasks and systematically avoid complex programmatic logic, resulting in inadequate code coverage. Currently, mitigating this premature termination requires continuous human-in-the-loop supervision. This heavy reliance on human intuition creates a bottleneck that negates the efficiency gains of automated generation. We propose SCATE, a framework for adaptive, automated supervision of coding agents that replaces human intervention during test generation. By formulating supervision as a contextual bandit problem, SCATE learns to select the most promising testing actions based on the current coverage and class testability metrics, maximizing coverage gains while minimizing wasted generation effort. Our empirical evaluation demonstrates that SCATE integrates seamlessly with different coding agents. When applied to GEMINI-CLI, it achieves 32.3% higher line coverage and 30.9% higher branch coverage than the agent-only baseline. A comparison with CLAUDE CODE confirms the framework dynamically adapts its policy to optimize each agent's unique strengths. SCATE also consistently outperforms state-of-the-art non-agentic approaches across all metrics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑