HarnessOpt-Bench:评估大型语言模型的 harness 优化能力
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
浏览论文内容
中文总结 AI 辅助
本研究提出HarnessOpt-Bench基准,评估前沿LLM在高成本随机评估下的端到端harness优化能力,发现优化器模型差异大于编码harness、原生harness非始终更优,该能力具可测性与提升空间。
中文摘要 AI 辅助
随着大型语言模型(LLM)越来越多地部署在智能体系统中,其能力不仅取决于模型权重,还取决于harness——围绕模型的提示词、工具、控制流、内存和编排代码。这使得自动化harness优化(即AI系统通过迭代和评估引导的方式改进harness)成为提升AI系统的重要途径,也是AI系统自身需要具备的高要求能力。然而,学界缺乏通用协议来评估前沿LLM在该任务上的表现。本文提出HarnessOpt-Bench,这是一个在高成本且随机评估场景下进行端到端harness优化的基准。优化器是一个结合了编码harness的LLM,它会接收目标智能体的种子harness、分级评估反馈以及固定的目标评估预算,然后编辑harness并提名最终候选方案,该方案的得分基于其在保留的测试分区上相对于种子harness的归一化增益,而该测试分区在整个搜索过程中保持不可访问。可信执行环境会强制执行评估边界、计量目标智能体的资源使用情况,并保留候选版本以供审计。我们将5个前沿LLM作为优化器,在共享编码harness和它们的原生harness下,针对4个下游任务进行了超过111次评分运行的评估。实验结果显示,优化器模型之间的差异大于它们所作用的编码harness之间的差异,原生harness并不始终更优,且增益在不同任务和种子方案之间存在显著差异。这些结果表明,harness优化是一种可测量且具有区分度的能力,存在巨大的提升空间。
英文摘要
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guided improvement of a harness by an AI system -- both an important route to improving AI systems and a demanding capability for AI systems themselves. Yet the community lacks a common protocol for measuring how well frontier LLMs perform at this task. We introduce HarnessOpt-Bench, a benchmark for end-to-end harness optimization under expensive and stochastic evaluation. An optimizer, an LLM paired with a coding harness, receives a target agent's seed harness, graded evaluation feedback, and a fixed target-evaluation budget. It edits the harness and nominates a final candidate, which is scored by its normalized gain over the seed on a held-out test partition that remains inaccessible throughout search. A trusted execution environment enforces the evaluation boundary, meters target-agent resource use, and preserves candidate versions for audit. We evaluate 5 frontier LLMs as optimizers both under a shared coding harness and under their native harnesses across 4 downstream tasks, over 111 scored runs. Experiment results show that optimizer models separate more than the coding harnesses they act through, native harnesses are not consistently superior, and gains vary substantially across tasks and seed regimes. These results establish harness optimization as a measurable and discriminative capability with large space for improvement.
发表机构
- Scale AI
机构由 AI 辅助整理,请以论文原文为准。