arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Evo-Harness:面向自进化智能体的上下文到 harness 的技能编译

Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents

Tianxin Wei, Zhan Shi, Minhua Lin, Bing He, Zewen Liu, Yisi Sang, Yuanchen Bei, Xuying Ning, Jiaru Zou, Ting-Wei Li, Xiao Lin, Yanjun Zhao, Chi Wang, Benoit Dumoulin, Dakuo Wang, Jingrui He, Hanqing Lu

arXiv 2608.15071首次发表:更新:

AI 中文总结

针对现有自进化 LLM 智能体方法的不足,提出 Evo-Harness 框架,通过上下文到 harness 的技能编译提炼单次执行的噪声上下文,在五个现实基准上验证了其有效性,为 LLM 智能体即时学习提供了原则性理解。

AI 中文摘要

从经验中学习对于开发有能力、可自我改进的大语言模型(LLM)智能体至关重要。现有方法通常通过反思、记忆、规则或技能从累积的轨迹中提取知识。然而,现实环境中的智能体会持续遇到新任务,通常仅提供一次改进机会。这些执行过程会产生丰富但噪声极高的上下文,将广泛有用的经验与特定任务的人工制品交织在一起。关键的是,现有工作很少在复杂的现实世界任务上验证其有效性,也很少分离出改进的潜在驱动因素。为解决这些差距,我们提出了在线 harness 学习,其中冻结的智能体通过在连续任务中不断更新结构化的 harness 来改进。该公式使我们能够通过提出的 Evo-Harness 系统研究关键的自我改进因素。其核心是上下文到 harness 的技能编译,将噪声大的单次执行提炼为可重复使用的技能 harness,用于跨域和主题级别的适应。为展示单次技能编译的功效,我们在五个现实基准(TerminalBench2、SWE-bench、CL-Bench、-bench、WebArena-Infinity)上进行了评估。我们的广泛分析证明了 Evo-Harness 的有效性,并为 LLM 智能体如何有效进行即时学习提供了原则性理解。我们的代码可在该 https URL 获取。

英文摘要

Learning from experience is critical for developing capable, self-improving large language model (LLM) agents. Existing methods typically extract knowledge from accumulated trajectories via reflection, memory, rules, or skills. However, agents in realistic environments continuously encounter novel tasks, often offering only a one-shot opportunity to improve. These executions yield rich but highly noisy contexts, entangling broadly useful lessons with task-specific artifacts. Critically, prior works rarely validate their effectiveness on complex real-world tasks or isolate the underlying drivers of improvement. To address these gaps, we formulate online harness learning, where a frozen agent improves by continually updating a structured harness across sequential tasks. This formulation enables a systematic study of key self-improvement factors through our proposed Evo-Harness. At its core, context-to-harness skill compilation distills noisy, single-shot executions into reusable skill harnesses for cross-domain and topic-level adaptation. To demonstrate the efficacy of one-shot skill compilation, we evaluate across five realistic benchmarks (TerminalBench2, SWE-bench, CL-Bench, -bench, WebArena-Infinity). Our extensive analysis demonstrates the effectiveness of Evo-Harness and provides a principled understanding of how LLM agents can effectively learn on the fly. Our code is available at https://github.com/A-EVO-Lab/a-evolve/tree/release/evo-harness.

CommentsEMNLP 2026 Main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑