arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ServeLearnBench:智能体如何从服务经验中自我改进?

ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?

Haizhong Zheng, Yizhuo Di, Ranajoy Sadhukhan, Shuowei Jin, Beidi Chen

arXiv 2610.07792首次发表:更新:

发表机构

Carnegie Mellon University; Amazon(卡内基梅隆大学; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ServeLearnBench通过演化环境流式数据集(EESD)评估智能体从服务经验中自我改进的能力,涵盖53个环境窗口和7,718个任务,发现任务能力与经验学习间存在差距、持续适应成本高且探索不足是关键瓶颈。

AI 中文摘要

大型语言模型智能体越来越多地被部署在真实环境中执行复杂任务。然而,在这些环境中正确行为所需的知识往往是隐性的、未公开的,并且会随时间变化。近期的持续学习框架旨在通过使智能体能够从服务经验中改进来应对这一挑战。然而,这些方法的有效性和局限性尚未得到充分刻画。现有基准仅提供部分覆盖:有些明确提供目标知识,有些假设静态环境,而那些支持持续适应的基准在规模上有限且知识多样性不足。为了支持系统性评估,我们形式化了一个演化环境流式数据集(EESD),其中智能体必须在隐藏策略演化时从交互和结果反馈中推断、应用和修正潜在环境知识,并引入了ServeLearnBench,涵盖零售支持、银行和销售话术生成,包含53个环境窗口和7,718个任务。我们评估了五种学习框架(RAG、Mem0、SkillOpt、Continual Harness和Prime),跨越六种模型(GPT-5.6 Terra、Opus 5、Kimi K3、GLM-5.3、DeepSeek V4.1 Flash和GLM-5.3 Flash),覆盖28个模型-框架组合和252次学习运行。我们的评估揭示了三个主要发现:任务能力与从经验中学习之间仍存在显著差距;持续适应成本高昂且可能降低已正确行为的性能;探索不足成为有效适应的关键瓶颈。总体而言,ServeLearnBench提供了一个受控测试平台,用于诊断这些局限性并追踪智能体通过服务经验持续可靠改进的进展。

英文摘要

Large language model agents are increasingly deployed to perform complex tasks in real-world environments. However, the knowledge required for correct behavior in these environments is often implicit, undisclosed, and subject to change over time. Recent continual-learning harnesses seek to address this challenge by enabling agents to improve from serving experience. Yet the effectiveness and limitations of these methods are not yet well characterized. Existing benchmarks provide only partial coverage: some explicitly provide the target knowledge, others assume a static environment, and those that support continual adaptation remain limited in scale and knowledge diversity. To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks. We evaluate five learning harnesses (RAG, Mem0, SkillOpt, Continual Harness, and Prime) across six models (GPT-5.6 Terra, Opus 5, Kimi K3, GLM-5.3, DeepSeek V4.1 Flash, and GLM-5.3 Flash), covering 28 model-harness pairs and 252 learning runs. Our evaluation reveals three main findings: a substantial gap remains between task capability and learning from experience; continual adaptation is costly and can degrade already-correct behavior; and insufficient exploration emerges as a key bottleneck to effective adaptation. Overall, ServeLearnBench provides a controlled testbed for diagnosing these limitations and tracking progress toward agents that continually and reliably improve through serving experience.

Comments28 pages, 13 figures. Code: https://github.com/Infini-AI-Lab/ServeLearnBench. Project website: https://infini-ai-lab.github.io/ServeLearnBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑