arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

正确的代码,有缺陷的贡献?SWE-CC:针对编码代理的仓库策略合规性基准测试

Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents

Hai Dang Truong, Rayner Goh, Thanh Le-Cong, Yintong Huo

arXiv 2610.06193首次发表:更新:

发表机构

Singapore Management University; Singapore University of Technology and Design(新加坡管理大学; 新加坡科技设计大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对编码代理的仓库策略合规性,提出SWE-CC基准,通过823条原子策略和全面审计机制评估500个任务,发现功能正确但合规性差,43.1%策略被违反。

AI 中文摘要

自主编码代理现在能够解决相当一部分真实的GitHub问题。然而,通过功能测试与产生可合并的高质量贡献在本质上有所不同。成熟的开源项目会发布针对仓库的贡献策略,涵盖风格、git、测试工作流等方面,以确保代码质量和长期可维护性。由于现有基准仅基于单元测试评估补丁,代理对仓库治理的合规性仍然未知。在本文中,我们引入了SWE-CC,一个评估自主软件工程中代码和过程合规性的基准。我们开发了一个半自动化的流水线,将12个开源仓库中的开发者文档转换为823条机器可检查的原子策略。SWE-CC引入了两个特性:1)轻量级、确定性的检查器函数,代表每条策略;2)一个全面的审计机制,检查代理的运行时行为和最终交付物。我们在从SWE-bench Verified扩展而来的500个端到端软件贡献任务中评估了代理工作流的合规性。我们在两种代理框架下对四个LLM的评估显示,现代代理存在编码合规性问题:尽管代理生成了功能正确的补丁,它们仍然违反了43.1%的适用项目策略,且近一半的违规发生在中间执行步骤中。这些结果表明,功能正确性并不能保证现实世界的就绪性,强调未来的软件工程代理必须可靠地符合仓库治理,以实现安全可信的部署。

英文摘要

Autonomous coding agents now resolve a substantial share of real-world GitHub issues. However, passing functional tests differs fundamentally from producing a high-quality contribution acceptable for merging. Mature open-source projects publish repository-specific contribution policies, spanning style, git, testing workflows, to ensure code quality and long-term maintainability. Because existing benchmarks evaluate patches solely on unit tests, agent compliance with repository governance remains unknown. In this paper, we introduce SWE-CC, a benchmark evaluating code and process compliance in autonomous software engineering. We develop a semi-automated pipeline that converts developer documentation across 12 open-source repositories into 823 machine-checkable atomic policies. SWE-CC introduces two features: 1) lightweight, deterministic checker functions that represent each policy, 2) a comprehensive auditing mechanism that inspects both agent runtime behaviors and final deliverables. We evaluate the compliance of agent workflows in 500 end-to-end software contribution tasks extended from SWE-bench Verified. Our evaluation of four LLMs under two agent scaffolds shows that modern agents suffer from coding compliance issues: although agents produce functionally correct patches, they still violate 43.1 percent of applicable project policies, with nearly half of all violations occurring during intermediate execution steps. These results show that functional correctness does not guarantee real-world readiness, highlighting that future software engineering agents must reliably conform to repository governance to enable safe and trustworthy deployment.

Comments39 pages, 6 figures. Benchmark and source code: https://github.com/dangtruong01/swe-cc-arxiv

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑