arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PonyEval:评估基于LLM的针对能力安全与Actor导向的Pony软件的程序修复

PonyEval: Evaluating LLM-Based Program Repair for Capability-Safe and Actor-Oriented Pony Software

Bang Xie, Hao Liu, Zhenyu Shi, Zhiyuan Peng, Xin Yin, Chenhao Ying, Yuan Luo, Haiming Jin, Wei Chen, Senjian Zhang, Shaocong Long

arXiv 2609.27832首次发表:更新:

AI 中文总结

针对Pony语言生态,构建了包含291个真实GitHub问题-拉取请求对的SWE-bench风格基准PonyEval,并通过严格审计与多模型评估,揭示生成补丁的条件解决率在10.21%至24.68%之间,为能力安全与actor导向软件的程序修复提供了可靠评测。

AI 中文摘要

仓库级问题解决基准已使可执行评估成为软件工程智能体的核心,但其语言覆盖仍集中在主流生态系统中。Pony呈现了一种不同的模式:它结合了actor、引用能力、提前编译以及快速演化的历史工具链,使得补丁生成和忠实回放都变得困难。我们引入了PonyEval,一个SWE-bench风格的基准,包含来自15个Pony仓库的291个真实GitHub问题-拉取请求对。每个实例绑定了一个问题陈述、一个历史基础提交、一个开发者黄金补丁、一个黑盒测试补丁以及一个可复现的运行时映射。冻结版本通过了一项离线审计,要求特定于问题的测试在基础状态上失败,并在黄金补丁后通过;它不包含重复的实例标识符或规范的仓库-PR对。在另一项全量集语义选择审计中,三个隔离的机器审阅者将所有291个实例标记为包含或排除;它们的Fleiss' kappa为0.8968,其中249个一致包含和32个一致排除。为了回放十一年的仓库历史,我们重建了覆盖289个唯一基础提交的72个运行时镜像,并在五个异构计算节点上验证了它们的可用性。我们定义了一个匹配评估,使用mini-SWE-agent 2.4.6对GPT-5.6-sol、DeepSeek-V4-Pro、GLM-5.2、MiniMax-M3和Kimi-K3,随后进行严格的补丁应用、编译和隐藏测试验证。在每种模型实际生成的补丁中,条件解决率范围从10.21%到24.68%。这些比率表征了生成补丁的质量,而非在所有291个基准任务上的成功。

英文摘要

Repository-level issue-resolution benchmarks have made executable evaluation central to software-engineering agents, but their language coverage remains concentrated in mainstream ecosystems. Pony presents a different regime: it combines actors, reference capabilities, ahead-of-time compilation, and a rapidly evolving historical toolchain, making both patch generation and faithful replay difficult. We introduce PonyEval, a SWE-bench-style benchmark of 291 real GitHub issue-pull-request pairs from 15 Pony repositories. Every instance binds an issue statement, a historical base commit, a developer gold patch, a black-box test patch, and a reproducible runtime mapping. The frozen release passes an offline audit requiring the issue-specific test to fail on the base state and pass after the gold patch; it contains no duplicate instance identifiers or canonical repository-PR pairs. In a separate full-set semantic selection audit, three isolated machine reviewers label all 291 instances as include or exclude; their Fleiss' kappa is 0.8968, with 249 unanimous inclusions and 32 unanimous exclusions. To replay eleven years of repository history, we reconstruct 72 runtime images covering 289 unique base commits and verify their availability on five heterogeneous compute nodes. We define a matched evaluation with mini-SWE-agent 2.4.6 for GPT-5.6-sol, DeepSeek-V4-Pro, GLM-5.2, MiniMax-M3, and Kimi-K3, followed by strict patch application, compilation, and hidden-test validation. Across the patches actually produced by each model, conditional resolution rates range from 10.21% to 24.68%. These rates characterize the quality of generated patches rather than success over all 291 benchmark tasks.

Comments9 pages, 1 figure, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑