发表机构
Queen’s University; Concordia University(女王大学; 康考迪亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对软件工程智能体回归测试成本高的问题,提出基于轨迹嵌入的确定性基准子集选择方法,在10%子集下中位误差低于5%,令牌成本降低约90%。
AI 中文摘要
自主软件工程智能体(SWE-agents)自动化编码任务。每次智能体更新可能需要重新运行完整基准以检测回归和性能改进,每次运行的成本高达数亿个LLM令牌,这使得评估成为瓶颈。一种解决方案是仅评估基准实例的子集。然而,简单的方法,如随机抽样或基于过去通过/失败结果的分层随机抽样,可能产生高方差和不具代表性的子集。我们转向智能体轨迹,即智能体所采取行动的逐步记录。我们提出了一种轨迹感知的子集选择方法,用基于轨迹嵌入的确定性选择取代随机抽样。我们首先根据最近一次完整测试运行中的测试结果对测试集实例进行分组,以保留历史通过/失败率,然后使用轨迹的嵌入空间选择子集。我们评估了76种子集选择配置,包括随机抽样、基于嵌入的选择、基于聚类的选择以及混合的候选名单后子抽样策略,涵盖三种回归场景:相同配置的重新运行、模型和配置更改以及智能体框架更改。我们最佳的轨迹感知方法是选择嵌入空间中距离每个结果组质心最近的基准实例。在我们评估的所有方法中,它实现了最低的估计误差。例如,当在选定的5%或10%测试实例子集上评估给定智能体版本时,我们的方法相对于典型抽样将平均估计误差降低了3-11%,相对于最强基线的95百分位抽样将最坏情况误差降低了4-11%,相对于典型抽样降低了38-46%。我们的结果表明,10%的轨迹感知子集将中位估计误差保持在5%以下,同时将令牌成本削减约90%。
英文摘要
Autonomous software engineering agents (SWE-agents) automate coding tasks. Each agent update may require re-running the full benchmark to detect regressions and improvements, at a cost of hundreds of millions of LLM tokens per run, which makes evaluation a bottleneck. One solution is to evaluate only a subset of benchmark instances. Yet, simple approaches, such as random sampling or stratified random sampling based on past pass/fail outcomes, risk producing high variance and unrepresentative subsets. We turn to agent trajectories, the step-by-step record of the actions an agent took. We propose a trajectory-aware subset selection approach that replaces random sampling with deterministic selection based on trajectory embeddings. We first group test set instances by their test outcome in a recent full test run to preserve the historical pass/fail rate, then select the subset using the trajectory's embedding space. We evaluate 76 subset selection configurations, including random sampling, embedding-based selection, clustering-based selection, and hybrid shortlist-then-subsample strategies, across three regression scenarios: same-configuration reruns, model and configuration changes, and agent framework changes. Our best trajectory-aware method is the one selecting benchmark instances closest to the centroid of each outcome group in the embedding space. It achieves the lowest estimation error among all methods we evaluate. For instance, when evaluating a given agent version on a selected subset of 5% or 10% of the test instances, our approach reduces the average estimation error by 3--11% and the worst-case error by 4--11% relative to the typical draw and 38--46% relative to the 95th-percentile draw of the strongest baseline. Our results show that a 10% trajectory-aware subset keeps the median estimation error below 5% while cutting token cost by roughly 90%.