发表机构
Conn Castle Studios(康恩城堡工作室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DeltaSelect提出一种开源方法,通过相关性和回归选择低成本任务集,用于编码智能体的A/B测试,在预算内实现高效且可靠的版本比较。
AI 中文摘要
编码智能体基准测试旨在进行广泛而全面的比较,而非频繁的开发决策。单次运行结果存在差异,完整测试套件成本高昂,且基准测试框架可能与实际使用的框架不同。在对DeepSWE已发布试验的重采样分析中,仅19.5%的任务(113项中的22项)与完整基准性能的第五百分位Pearson相关系数达到至少0.50。本文提出DeltaSelect,一种开源方法,该方法利用Pearson相关识别单次运行结果能持续跟踪完整基准性能的任务,通过线性回归将分数验证器结果映射为通用分数,并在美元预算内选择固定的任务集。DeltaSelect旨在用于开发过程中重复的基线版本与候选版本比较,而非模型排名。在gpt-5.6-luna低推理能力的案例研究中,DeltaSelect被用于修订自定义技能和指令。在13次评估中,按2026年8月16日发布的价格计算,记录成本为27.86美元。采用版本的成本比初始版本低58.1%(1.75美元对比4.18美元;p=0.008),而校准分数更高(42.36%对比36.46%;发布模拟方差p=0.326)。
英文摘要
Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.50 with full-benchmark performance. The paper presents DeltaSelect, an open-source method that identifies tasks whose one-run results consistently track full-benchmark performance using Pearson correlation, maps fractional verifier results to a common score using linear regression, and selects a fixed task set within a dollar budget. DeltaSelect is intended for repeated baseline-versus-candidate comparisons during development, not model rankings. In a gpt-5.6-luna low-reasoning case study, DeltaSelect was used to revise custom skills and instructions. Across 13 evaluations, the recorded cost was USD 27.86 at rates published August 16, 2026. The adopted version cost 58.1% less than the initial version (USD 1.75 versus USD 4.18; p=0.008), while the calibrated score was higher (42.36% versus 36.46%; published-analog variance p=0.326).