发表机构
Zhejiang University; Alibaba Group; City University of Hong Kong; The Hong Kong University of Science and Technology; University of Michigan(浙江大学; 阿里巴巴集团; 香港城市大学; 香港科技大学; 密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出自演化智能体harness Living-Harness,基于Evolution-SOP生成两类过程性知识,在8个交互式环境中大幅提升Pass@1,支持跨模型主干复用演化后的harness状态。
AI 中文摘要
大型语言模型(LLM)智能体可在单轮内或重试后从故障中恢复,但相同的执行故障会在后续任务中复发,因为单轮后的反馈极少会修改指导未来交互的持久化harness。静态harness通过固定工具、上下文、记忆和工作流结构提升可靠性,但部署后保持不变。我们提出Living-Harness,一种自演化智能体harness,它将每个完成的轨迹及其评估器信号转化为有界harness更新的后验证据。在领域级Evolution-SOP(标准操作流程)的指导下,Living-Harness提取单轮抽象和结构化更新证据,并写入两种互补形式的过程性知识:记录触发条件、故障模式和恢复动作的单轮记忆,以及记录状态节点、修复边和转移规则的状态图。更新后的harness状态会被检索以指导未来交互,而工具和基础上下文保持冻结,使过程性修复能在演化周期中累积。在从τ²-Bench和MultiWOZ-2.4衍生的8个交互式环境中,Living-Harness将平均Pass@1分别比最强的交互式基线提升10.07和9.91个百分点,且支持在不同模型主干上仅检索复用演化后的harness状态。
英文摘要
Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedback rarely revises the persistent harness that guides future interactions. Static harnesses improve reliability through fixed tools, context, memory, and workflow structures, but remain unchanged after deployment. We propose $\textbf{Living-Harness}$, a self-evolving agent harness that converts each completed trajectory and its evaluator signals into posterior evidence for bounded harness updates. Guided by a domain-level $\textbf{Evolution-SOP}$ ($\textbf{S}$tandard $\textbf{O}$perating $\textbf{P}$rocedure), Living-Harness extracts an episode abstraction and structured update evidence, and writes two complementary forms of procedural knowledge: episodic memory that records trigger conditions, failure patterns, and recovery actions, and a state graph that records state nodes, repair edges, and transition rules. The updated harness state is retrieved to guide future interactions, while tools and base context remain frozen, allowing procedural repairs to accumulate across evolution cycles. On eight interactive environments derived from $τ^2$-Bench and MultiWOZ-2.4, Living-Harness improves average Pass@1 over the strongest interactive baseline by 10.07 and 9.91 percentage points, respectively, and supports retrieval-only reuse of the evolved harness state across model backbones. Our code will be made publicly available soon at https://github.com/anotherbricki/Living-Harness.