arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

什么因素导致软件问题解决任务对智能体而言难度较高?

What Makes Software Issue Resolution Tasks Difficult for Agents?

Ebtesam Al-Haque, Brittany Johnson

arXiv 2608.18280首次发表:更新:

发表机构

George Mason University(乔治梅森大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出测量框架,基于CoderForge-Preview数据集的实证分析,发现软件问题解决任务难度可通过静态特征预测,核心驱动因素为补丁碎片化与代码库规模,为构建难度可控的智能体评估基准奠定基础。

AI 中文摘要

背景:智能体系统的进展正迅速在基准测试中达到饱和状态。尽管这一现象常被讨论,但由于缺乏对任务难度的控制和表征,基准测试分数仍难以解读。更具体地说,目前我们几乎不理解是什么让一项任务比另一项更难,以及在多大程度上可以通过静态任务属性预测任务难度。目标:我们提出一种测量框架,以调查并系统量化软件任务的哪些结构属性对应智能体在问题解决任务中的成功率。方法:我们对CoderForge-Preview(迄今为止最大的编码智能体轨迹公开数据集)进行了大规模实证研究,从任务补丁、代码库和提示中提取特征。我们使用集成方法、SHAP归因和效应量分析评估每个特征对任务结果的预测能力。结果:我们发现任务难度可通过静态特征较好地预测(AUC=0.863),且主要由补丁碎片化程度和代码库规模驱动。提示语言特征在中等难度任务的顶级贡献者中显现,揭示了难度的分层结构。结论:问题解决任务的难度编码在其结构中,这支持了静态、事前的难度估计,并为构建难度可控的智能体评估基准测试奠定了基础。

英文摘要

Background. Advances in agentic systems are simultaneously, and rapidly, saturating benchmarks. Despite this often discussed phenomena, benchmark scores remain difficult to interpret due to the lack of control and characterization of task difficulty. More specifically, we currently have little understanding of what makes one task harder than another, and to what extent task difficulty is predictable from static task properties. Aims. We propose a measurement framework to investigate and systematically quantify what structural properties of software tasks correspond to agent success rates for issue resolution tasks. Method. We conducted a large scale empirical study on CoderForge-Preview, the largest open dataset of coding agent trajectories to date, by extracting features across task patch, repository and prompt. We evaluated the predictive power of each feature against task outcomes using ensemble methods, SHAP attribution, and effect size analysis. Results We found that task difficulty is substantially predictable from static features (AU C = 0.863) and is largely driven by patch fragmentation and repository scale. Prompt linguistic features become visible among top contributors for tasks in the mid-band, revealing a layered structure of difficulty. Conclusion. The difficulty of an issue resolution task is encoded in its structure. This enables static, pre-hoc difficulty estimation and lays the groundwork for difficulty-controlled benchmark construction for evaluation of agents.

CommentsTo appear in ESEM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑