发表机构
Allen Institute for AI; University of Washington(艾伦人工智能研究所; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究重新审视大语言模型智能体的自动 harness 进化评估,通过在可比条件下与基线比较及在保留任务上测试,发现其不总优于简单方法且泛化有限,对其有效性提出质疑,强调需更公平评估协议和基准。
AI 中文摘要
我们重新审视了大语言模型智能体的自动 harness 进化评估。现有 harness 进化方法使用单元测试用例来搜索 harness 配置,然后在相同的公共基准上报告最终性能。此协议引发了两个基本问题。首先,harness 进化本身是一个迭代搜索过程,应在匹配的反馈和推理预算下与简单的任务级搜索基线进行比较。其次,由于搜索和最终评估共享相同的基准,报告的收益可能会过度拟合该特定任务集。为解决这些问题,我们进行了广泛评估,在可比的反馈和推理预算下将 harness 进化与简单的测试时缩放和发现基线进行比较,并在保留任务上评估进化后的 harness 以评估发现的改进是否具有通用性。在使用 GPT - 5.4 和 Claude Opus 4.6 对 Terminal - Bench 2.1 进行的实验表明,自动 harness 进化并不总是优于简单的测试时缩放方法,并且泛化能力有限。我们的结果对自动 harness 进化的有效性提出了重要问题,并强调了对自动 harness 设计采用更公平评估协议和基准的必要性。
英文摘要
Harness evolution is an iterative search procedure that repeatedly evaluates and revises candidate harnesses used for LLM agents using task feedback. We revisit the evaluation of such automatic harness evolution procedures and identify two fundamental issues in the protocol. First, prior work does not compare these approaches with simple task-level search baselines under matched feedback and inference budgets. Second, prior work searches for harness configurations using verification signals (e.g., unit test cases) drawn from the same benchmarks on which it reports the final performance of the evolved harnesses, violating the standard separation between training and test data. To address this, we compare automatic harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and evaluate evolved harnesses on held-out tasks to assess generalization. Following prior work, we experiment on Terminal-Bench 2.1 and find that automatic harness evolution fails to outperform simple test-time scaling methods both with and without test cases, and exhibits limited generalization. However, we find that long-horizon games are a promising setting for automatic harness evolution, as they are difficult enough to leave headroom, rely heavily on adaptation to out-of-distribution dynamics, and provide granular feedback by design. In these settings, task-specific harness evolution improves over the search baseline by 80% on ARC-AGI-3 and by 11% on EdgeBench games under matched budgets. Together, these findings highlight the need for matched-budget baselines and held-out evaluation to distinguish genuine harness improvements from benchmark-specific search and overfitting, and point to a more careful characterization of when automatic harness evolution is actually useful. Our code is available at https://github.com/rethinking-harness-evolution.