相同运行,不同结果:对开放权重模型上的AI编码智能体进行基准测试
Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
浏览论文内容
中文总结 AI 辅助
本研究通过584次重复实验,发现AI编码智能体在相同配置下运行结果差异大,少数运行排名不可靠,建议将智能体与模型作为配对多次评估,并优先拒绝违规运行以提升交付模型质量。
中文摘要 AI 辅助
已知同一编码智能体的重复运行会给出不同的基准分数。我们通过在一个机器学习任务上的密集复现,探究这种变异性对团队在自身任务上运行智能体意味着什么:一个智能体改进用于航空延误的XGBoost分类器的训练代码,而一个它从未见过的保留集对结果进行评分。在584次运行中,我们比较了六个智能体在六个开放权重模型端点上的表现,在固定设置下将六个智能体-模型配对各运行52次,并对其中的三个配对在同一家族的更大模型上重复运行。同一配对的相同运行之间的差异大于不同配对之间的差异,因此少数几次运行的比较对其排名不可靠;要解决我们观察到的智能体差异,每个智能体需要数十次到一百多次运行。在更大模型上的运行得分明显更高,但差距小于一次运行间的标准差,且这一差距在某个智能体上比其他智能体大两倍以上。不到二十分之一的运行违反了任务的数据规则,但这些运行却获得了最高分。首先拒绝这些运行,并在几次尝试中保留最佳合规结果,可靠地改进了交付的模型,尽管少数几次运行无法对智能体进行排名。在后来年份的航班上,交付的模型仅保留了相对初始代码三分之一的增益。按标价计算,同一模型上两个智能体之间的成本差异超过二十倍,主要源于提示缓存。智能体和模型应作为配对进行评估,通过重复尝试,并在质量旁边报告合规性。数据、代码和每个交付的程序:此https URL
英文摘要
Repeated runs of the same coding agent are known to give different benchmark scores. We ask what that variation means for a team running an agent on its own task, by intensive replication on one machine-learning task: an agent improves the training code of an XGBoost classifier for airline delays, and a holdout it never sees scores the result. Across 584 runs, we compare six agents on six open-weight model endpoints, run six agent-model pairings 52 times each under fixed settings, and repeat three of them on a larger model from the same family. Identical runs of one pairing varied more than the pairings differed from one another, so comparisons of a few runs ranked them unreliably; resolving the agent differences we observed would take tens to more than a hundred runs of each. Runs on the larger model scored clearly higher, but by less than one run-to-run standard deviation, and the gap was more than twice as large with one agent as with the others. Fewer than one run in twenty broke the task's data rules, but those runs held the highest scores. Rejecting those runs first and keeping the best compliant result among a few attempts reliably improved the delivered model, even though a few runs could not rank the agents. On flights from a later year, the delivered models kept only a third of their gain over the starting code. At list prices, cost differed more than twentyfold between two agents on the same model, mostly through the prompt cache. Agents and models should be evaluated as pairings, over repeated attempts, with compliance reported beside quality. Data, code and every delivered program: https://github.com/earino/identical-runs-different-results
发表机构
- Central European University(中欧大学)
- Epoch
机构由 AI 辅助整理,请以论文原文为准。