arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28587cs.SEcs.AI

PAIChecker:揭示与检查类SWE-Bench基准中的PR-Issue不匹配问题

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

Manyi Wang, Junjielong Xu, Pinjia He

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对类SWE-Bench基准中13.6%的PR-Issue不匹配问题,提出多智能体系统PAIChecker,经实验验证其在两类基准上的二元准确率最高达92.12%

中文摘要 AI 辅助

类SWE-Bench基准被广泛用于评估大语言模型(LLM)的问题解决能力,其通常遵循通用构建流程:通过从PR(拉取请求)描述中提取问题引用,将PR与其关联问题配对;问题描述用作问题陈述,PR补丁作为测试预言。然而,由于大型仓库开发与维护的固有复杂性,此类PR-Issue配对在实际中常出现不匹配。本研究系统分析SWE-bench Verified实例,发现13.6%的实例存在不匹配,涉及5种模式及11种细粒度场景。为实现未来此类基准的可靠、可扩展构建,我们提出PAIChecker——一种用于检查类SWE-Bench基准中PR-Issue不匹配的多智能体系统。具体而言,PAIChecker采用三阶段设计,结合特定模式识别、跨智能体标签合成与代码级验证,从而实现更准确、可泛化且逐步验证的检测。在SWE-Gym和SWE-bench Multilingual上的实验表明,PAIChecker在所有四种LLM主干模型中均取得最佳性能,二元准确率分别高达92.12%和91.67%。

英文摘要

SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a common construction pipeline: each PR (Pull Request) is paired with its linked issue by extracting issue references from the PR description; the issue description is used as the problem statement, and the PR patch serves as the test oracle. However, due to the inherent complexity of developing and maintaining large repositories, such PR-Issue pairings are often misaligned in practice. In this work, we systematically study SWE-bench Verified instances, finding that 13.6% exhibit misalignment across five patterns in eleven fine-grained scenarios. To enable reliable and scalable construction of those benchmarks in the future, we propose PAIChecker, a multi-agent system for checking PR-Issue misalignment in SWE-bench-like benchmarks. Specifically, PAIChecker adopts a three-phase design that combines specific pattern identification, cross-agent label synthesis, and code-level validation, thereby enabling more accurate, generalizable, and progressively verified detection. Experiments on SWE-Gym and SWE-bench Multilingual show that PAIchecker achieves the best performance across all four LLM backbones, reaching up to 92.12% and 91.67% binary accuracy, respectively.

发表机构

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑