发表机构
Sun Yat-sen University; Huawei Cloud Computing Technologies Co., Ltd.; Zhejiang University(中山大学; 华为云计算技术有限公司; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对软件工程智能体基准评估成本高的问题,提出PTA-IRT框架融合过程与结果信号,在低预算下优于基线,提升了基准评估效率。
AI 中文摘要
在真实基准上评估软件工程智能体成本高昂,因为每个任务可能需要多步骤的代码探索、修改和测试执行。现有高效评估方法选择代表性子集来估计全基准性能,但大多仅关注结果:它们拟合历史通过/失败响应矩阵或静态任务语义,忽略智能体解决问题的过程。我们提出PTA-IRT,一种特权轨迹感知项目反应理论框架,融合过程与结果信号。历史执行轨迹提供超越通过/失败的过程级证据,如探索的上下文、尝试的编辑和解决路径,PTA-IRT将这些作为特权信息用于校准子集选择和能力估计。在校准预算较低时,PTA-IRT在四个软件工程基准上的分数和排名恢复方面始终优于先前的IRT基线。代码和数据可在指定URL公开获取。
英文摘要
Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-benchmark performance, but are largely result-only: they fit historical pass/fail response matrices or static task semantics, discarding how agents solve problems. We propose PTA-IRT, a Privileged Trajectory-Aware Item Response Theory framework that fuses process and outcome signals. Historical execution trajectories supply process-level evidence beyond pass/fail, such as explored context, attempted edits, and solving paths, which PTA-IRT uses as privileged information for calibration subset selection and ability estimation. Under low calibration budgets, PTA-IRT consistently outperforms prior IRT baselines on score and ranking recovery across four SWE benchmarks. Code and data are publicly available at https://github.com/DeepSoftwareAnalytics/PTA-IRT.
CommentsUnder review