arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HERA:工具环境协同进化实现可靠的智能体弃权(不执行)

HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention

Han Luo, Bingbing Wen, Guang Yang, Zora Zhiruo Wang, Pan Lu, Lucy Lu Wang

arXiv 2610.06563首次发表:更新:

发表机构

University of Washington; University of Leeds; Carnegie Mellon University; Stanford University; Allen Institute for AI(华盛顿大学; 利兹大学; 卡内基梅隆大学; 斯坦福大学; 艾伦人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出HERA框架,通过工具框架-环境协同进化自动生成可行与不可行任务对,提升智能体弃权(不执行)可靠性,将弃权准确率从61.7%提升至83.3%,并跨19个LLM平均提升15.3个百分点。

AI 中文摘要

大型语言模型(LLM)智能体在复杂的工具使用环境中行动能力日益增强,但它们往往无法识别任务不可行且不存在有效解决方案的情况。近期工作已将这一可靠性差距形式化为智能体弃权(不执行)问题,现有方法通常针对固定任务集优化模型或智能体工具框架,导致对未见故障模式的泛化能力有限。我们提出HERA,一个用于智能体弃权(不执行)的工具框架-环境协同进化框架。HERA包含:(i)一个自动构建可验证的可行与不可行任务对的流水线,通过应用受控环境变异,将可解决任务转化为需要弃权(不执行)的案例;(ii)一个协同进化过程,其中先前任务上的性能失败被用于驱动工具框架适应,并生成针对先前弱点的新执行环境和任务。在保留评估任务上,来自HERA的进化工具框架将弃权(不执行)准确率从61.7%提升至83.3%,同时将可行任务完成率从68.3%提升至76.7%,在所比较方法中实现了最高的弃权(不执行)准确率和可行任务完成率。所得最佳工具框架可迁移至其他19个LLM,无需任何模型特定优化即可将弃权(不执行)准确率平均提升15.3个百分点,并使较小模型以估计降低85%的成本达到更强大模型的性能。

英文摘要

Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments, yet they often fail to recognize when tasks are infeasible and no valid solution exists. Recent work has formalized this reliability gap as the problem of agentic abstention, and existing approaches typically optimize a model or agent harness against a fixed set of tasks, leading to limited generalization to unseen failure modes. We introduce HERA, a framework for harness-environment co-evolution for agentic abstention. HERA consists of (i) a pipeline to automatically construct verifiable pairs of feasible and infeasible tasks by applying controlled environment mutations that transform solvable tasks into cases requiring abstention, and (ii) a co-evolution procedure in which performance failures on previous tasks are used to drive harness adaptation and generate new execution environments and tasks geared towards previous weaknesses. On held-out evaluation tasks, an evolved harness from HERA improves abstention accuracy from 61.7% to 83.3% while improving feasible-task completion from 68.3% to 76.7%, achieving the highest abstention and feasible-task completion among the compared methods. The resulting best harness transfers across 19 other LLMs, improving abstention accuracy by 15.3 percentage points on average without any model-specific optimization, and enabling smaller models to match the performance of more powerful models at an estimated 85% lower cost.

Comments23 pages. Project page: https://hera-bench.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑