发表机构
Poisson Lab, Huawei(华为泊松实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TTSE提出双轨在线自进化框架,分离环境事实与任务程序,通过决策理论分析并实验验证,在多个基准上提升智能体任务适应性与性能。
AI 中文摘要
随着大型语言模型(LLM)智能体被应用于持续交互的环境中,驱动其自身能力的进化成为实现长期自主性的核心问题。目前,环境知识通常被视为外部固定输入,而非智能体持续进化的一部分。强化学习方法通常通过环境交互来优化策略,但往往仅适应固定的任务分布或单一环境。本文提出TTSE(双轨自进化),一种双轨在线自进化框架,将进化知识分为FACT(环境事实,其可靠性通过交互证据持续验证)和TIP(任务条件下的实施程序)。从决策理论的角度,我们将智能体的超额风险分解为环境表示遗憾和条件执行遗憾,刻画了环境条件策略严格优于条件无关策略的条件,并将下游风险以FACT识别误差和跨条件不匹配成本进行界定。在实践中,TTSE在GDPevo上的消融实验验证了双轨进化的优势。在经典智能体任务基准ALFWorld和ScienceWorld上,TTSE进一步展示了优越的任务适应性。此外,TTSE与现有技能自进化方法广泛兼容;结合Bayesian-Agent算法,单轨消融验证了双轨优势,在SOPBench的五个主要领域上,三次独立重复中显著提高了总分。最后,在真实端到端任务基准PinchBench上,TTSE通过基于检索的注入集成到通用智能体框架中,并在三次独立运行中稳定优于基线。
英文摘要
As Large Language Model (LLM) agents are applied in continuously interactive environments, driving the evolution of their own capabilities becomes a core problem for achieving long-term autonomy. Currently, environmental knowledge is typically treated as an external fixed input rather than as part of the agent's ongoing evolution. Reinforcement learning methods usually optimize policies through environmental interaction but tend to adapt only to fixed task distributions or single environments. This paper proposes TTSE (Two-Track Self-Evolution), a dual-track online self-evolution framework that separates evolving knowledge into FACT (environmental facts, whose reliability is continuously verified through interaction evidence) and TIP (task-conditioned implementation procedures). From a decision-theoretic perspective, we decompose the agent's excess risk into environment-representation regret and conditional-execution regret, characterize the conditions under which environment-conditioned policies strictly outperform condition-agnostic policies, and bound the downstream risk in terms of FACT identification error and cross-condition mismatch cost. In practice, TTSE's ablation experiments on GDPevo validate the advantage of dual-track evolution. On the classic agent task benchmarks ALFWorld and ScienceWorld, TTSE further demonstrates superior task adaptation. Moreover, TTSE is broadly compatible with existing skill self-evolution methods; combined with the Bayesian-Agent algorithm, a single-track ablation validates the dual-track advantage, substantially improving the aggregate score across the five major domains of SOPBench over three independent repetitions. Finally, on the real end-to-end task benchmark PinchBench, TTSE is integrated into a general agent framework via retrieval-based injection and stably outperforms the baseline across three independent runs.
Comments20 pages, 2 figures, 18 tables