发表机构
School of Data Science, Fudan University; College of Computer Science and Artificial Intelligence, Fudan University(复旦大学数据科学学院; 复旦大学计算机科学与人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体控制策略进化中验证成本高的问题,提出HarnessLens框架,通过行为感知的选择性验证,在提升保留任务性能的同时大幅降低评估预算。
AI 中文摘要
智能体控制策略(agent harness)决定了语言模型智能体如何使用指令、工具和运行时组件,但调整这些控制策略需要付出高昂的验证成本。现有的“提出并验证”方法通常会在固定任务集上对每个候选策略进行评分,在无关行为上浪费大量试错(rollout),且总评分会掩盖特定的性能退化。我们提出了HarnessLens,这是一种用于自动控制策略进化的预算感知框架。HarnessLens同时探索任务空间和用户可配置组件,从执行轨迹中推导候选修改方案,并使用归因证据门在与行为相关的任务上选择性验证每个候选方案。在三种智能体控制策略和四个基准测试中,HarnessLens将平均保留任务性能提升了7.6%-13.6%,同时消耗的评估预算显著少于对比基线。这些结果表明,具有显式归因的感知行为验证能够在有限的交互预算下实现更可靠、样本效率更高的控制策略进化。我们的代码可在此处获取:https://
英文摘要
Agent harnesses shape how language-model agents use instructions, tools, and runtime components, but adapting these harnesses requires costly verification. Existing propose-and-verify methods typically score every candidate on a fixed task set, wasting rollouts on unrelated behaviors and allowing aggregate scores to obscure specific regressions. We introduce HarnessLens, a budget-aware framework for automated harness evolution. HarnessLens jointly explores the task space and user-configurable components, derives candidate modifications from execution trajectories, and selectively verifies each candidate on behavior-relevant tasks using an attributable-evidence gate. Across three agent harnesses and four benchmarks, HarnessLens improves average held-out performance by 7.6-13.6% while consuming substantially less evaluation budget than competing baselines. These results demonstrate that behavior-aware verification with explicit attribution enables more reliable and sample-efficient harness evolution under constrained interaction budgets. Our code is available at https://github.com/jhxu5214/HarnessLens.
Comments17 pages, 6 figures