arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SWE-Prometheus:衡量真实仓库中的工程治理改进

SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

Jiajun Wu, Leixin Sun, Zihan Tan, Yitao Liu, Shuo Li, Jiaru Qian, Yuxin Wu, Shanghaoran Quan, Chuangxin Zhao, Yangxu Liao, Yang Liu, Bin Chong, Guancheng Wan

arXiv 2609.29465首次发表:更新:

发表机构

CosmosMind; Peking University; Tsinghua University; HKUST; ModCraft(CosmosMind; 北京大学; 清华大学; 香港科技大学; ModCraft)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SWE-Prometheus基准,用于评估编码智能体在真实仓库中改进工程治理的能力,通过多维度证据和教师评分衡量改进与行为保持,并揭示治理工件与执行支持改进的差异。

AI 中文摘要

基于大型语言模型的编码智能体在仓库级软件工程任务上取得了显著进展。然而,现有的仓库基准通常从人工识别的问题出发,评估补丁是否满足功能性信号。我们提出了SWE-Prometheus,一个面向更广泛任务——改进仓库工程治理——的基准。每项任务提供一个固定的快照和一个开放式目标,要求智能体识别风险、确定干预措施的优先级,并验证由此产生的变更。SWE-Prometheus通过配对证据、干净环境探针、行为门控以及两位独立教师对相同证据的评分,评估六个治理维度。该基准包含60个仓库;在共享的22个仓库公共子集上评估了十个模型,平均归一化治理改进(NGI)范围从0.0568到0.5760,观察到的行为破坏率范围从0%到23%。在冻结的十个仓库批次上,一个不依赖仓库的模板获得了平均NGI 0.272,但其收益集中在测试与持续集成、质量门禁和文档方面;在可复现环境和依赖与安全方面,它未能在任何仓库上实现改进。该基线使得区分“添加治理工件”与“产生可执行支持的改进”成为可测量的。无操作条件的NGI中位数为零,标准差为0.073;两位教师对相同的无操作证据在60个维度评分中有57个完全一致。对于两个条件均值最高的系统,共同有效的NGI相似,而包含行为失败的全池比较则偏向Kimi-K3。这些结果表明,仓库治理评估应同时报告改进、行为保持、证据质量和覆盖率。

英文摘要

Large language model based coding agents have made substantial progress on repository-level software engineering tasks. Existing repository benchmarks, however, usually start from a human-identified issue and evaluate whether a patch satisfies a functional signal. We present SWE-Prometheus, a benchmark for the broader task of improving repository engineering governance. Each task provides a fixed snapshot and an open-ended objective, requiring the agent to identify risks, prioritize interventions, and verify the resulting changes. SWE-Prometheus evaluates six governance dimensions through paired evidence, clean-environment probes, behavior gates, and two independent teacher ratings of the same evidence. The benchmark contains 60 repositories; ten models are evaluated on a shared 22-repository public subset, where mean Normalized Governance Improvement ranges from 0.0568 to 0.5760 and observed behavior-breakage rates range from 0% to 23%. On a frozen ten-repository batch, a repository-blind template obtains mean NGI 0.272, but its gains concentrate in Tests & CI, Quality Gates, and Documentation; it improves Reproducible Environment and Dependency & Security on none of the repositories. This baseline makes the distinction between adding governance artifacts and producing execution-backed improvements measurable. The no-op condition has median NGI zero and standard deviation 0.073; two teachers agree exactly on 57 of 60 dimension scores for the same no-op evidence. For the two highest conditional-mean systems, common-valid NGI is similar, while full-pool comparisons that include behavior failures favor Kimi-K3. These results show why repository-governance evaluation should report improvement, behavior preservation, evidence quality, and coverage together.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑