发表机构
University of Chinese Academy of Sciences; Microsoft; Institute of Automation, Chinese Academy of Sciences; Nanjing University(中国科学院大学; 微软公司; 中国科学院自动化研究所; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出AgentStream框架,评估三种流式场景下三种基础模型的五种自进化LLM智能体方法,发现自进化表现受场景、模型能力等影响,为方法选择提供指导。
AI 中文摘要
大语言模型(LLM)智能体可通过从自身积累的经验中持续改进实现自进化,但现有研究大多采用独立评估方式,导致对自进化智能体在现实流式场景(智能体需适应多样复杂任务流)中的行为认知不足。为解决这一问题,本文提出AgentStream,这是一个统一框架,通过将智能体基准组织为可配置任务流,在测试时实例化孤立(Isolated)、顺序(Sequential)和交错(Interleaved)三种流式场景(逐步改变任务流的范围与领域构成),以此评估涵盖不同进化组件的自进化智能体。在上述场景中,本文组合评估了三种前沿基础模型上的五种代表性自进化方法,解析了模型能力、方法架构与流式场景如何共同影响自进化。结果表明:自进化可靠性随流式场景变化,自进化的收益受模型能力限制且与模型强度呈非单调关系,不存在在所有模型和场景中均占优的单一方法。这些发现为跨模型和流式场景选择自进化方法提供了具体指导,总体而言,本文主张应在现实任务流而非孤立单任务设置下评估自进化智能体。
英文摘要
Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a unified framework that evaluates self-evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the \texttt{Isolated}, \texttt{Sequential}, and \texttt{Interleaved} streaming scenarios at test time, which progressively vary the scope and domain composition of the stream. Over these scenarios, we combinatorially evaluate five representative self-evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self-evolution. Our results show that self-evolution reliability varies across streaming scenarios, the benefit of self-evolution is gated by model capability and is non-monotonic in model strength, and no single method dominates across models and scenarios. These findings offer concrete guidance for selecting self-evolving methods across models and streaming scenarios. Overall, we advocate that self-evolving agents should be evaluated under realistic task streams rather than isolated single-task settings.
CommentsCode is available at https://github.com/Jasper-Yan/AgentStream