发表机构
Queen’s University(女王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出SNC轮廓表征编码任务需求,发现基准测试类别标签不可靠,智能体行为可揭示黄金方案未体现的需求,任务需求与成功相关且成功行为具家族特异性。
AI 中文摘要
智能体软件工程基准测试通常用“bug修复”或“功能实现”等名义类别标签来概括,但相同标签的基准测试却通过截然不同的整理流程构建。因此,标签几乎无法揭示基准测试所要求的软件工程工作。我们引入Spread–Novelty–Centrality(SNC)轮廓,这是基于实证软件工程研究构建的、针对仓库级编码任务需求的三轴表征方法。我们将该轮廓应用于5个广泛使用的基准测试,以及两个模型家族在三个规模下的14922条轨迹,并报告三项发现:(1)标签是任务需求的不可靠代理,因为每对基准测试在至少两个SNC轴上存在统计差异,且这种差异可追溯至特定的整理决策;(2)智能体行为能揭示人类编写的黄金解决方案无法体现的需求:当问题陈述未提供提示时,智能体生成的解决方案比黄金方案更大;当整理流程夸大黄金方案时,智能体生成的解决方案更小。任务的表述方式会影响智能体的产出;(3)任务需求与成功呈均匀相关,对于每个模型家族和规模,已解决的运行集中在低SNC区域,而成功的行为特征则是家族特有的:Claude通过匹配黄金方案的范围取得成功,其文件级奇偶性占比从最小规模的0.17升至最大规模的0.54;Qwen在所有规模下通过超出黄金方案的范围取得成功,而编辑过少则会导致两个家族均失败。
英文摘要
Agentic software engineering benchmarks are typically summarized by nominal category labels such as "bug fix" or "feature implementation," yet benchmarks carrying the same label are built through very different curation pipelines. A label thus reveals little about the engineering work a benchmark demands. We introduce the Spread--Novelty--Centrality (SNC) profile, a three-axis characterization of the demands of repository-level coding tasks, grounded in empirical software engineering research. We apply the profile to five widely used benchmarks and 14,922 trajectories of two model families at three scales, and report three findings. (1) A label is an unreliable proxy for task demands, as every pair of benchmarks is statistically separated on at least two SNC axes, and the separations trace back to specific curation decisions. (2) Agent behaviour reveals demands that the human-written gold solution cannot. Agents produce larger solutions than the gold where problem statements withhold hints and smaller ones where curation inflates the gold. How a task is phrased shapes what an agent produces. (3) Task demands correlate with success uniformly, with resolved runs concentrating in the low-SNC region for every family and scale, whereas the behavioural signatures of success are family-specific. Claude succeeds by matching the scope of the gold solution, and its parity share on files rises from $0.17$ at the smallest scale to $0.54$ at the largest. Qwen succeeds by exceeding the gold scope at every scale, and editing too little marks failure for both families.