arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EASE:面向自进化智能体的行为自适应技能策展

EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents

Zhen Xiong, Qiaoyu Tan

arXiv 2609.36746首次发表:更新:

发表机构

New York University; New York University Shanghai(纽约大学; 上海纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EASE提出行为自适应技能策展框架,通过在线行为画像和强化学习训练单一策展器,适应不同执行器行为,在多个基准上超越基线,减少技能数量并提升检索与编辑效用。

AI 中文摘要

智能体技能为自进化智能体提供了一种轻量级机制,无需更新模型参数即可积累可复用的程序性知识。然而,现有的学习型技能策展器通常在优化策展时未显式建模下游执行器的行为。我们证明,这可能导致系统性的跨执行器性能退化:使用不同执行器训练的策展器在与各自训练执行器配对时表现最佳,这表明有效的技能策展依赖于执行器。我们形式化了行为自适应技能策展,并引入EASE框架,该框架学习一个单一的策展器,使其决策适应不同执行器的行为。EASE维护一个近期执行模式的在线行为画像,并将策展器基于该画像、当前轨迹和检索到的技能进行条件化,以从不断演化的技能库中添加、修改或移除技能。我们使用强化学习在多个冻结执行器上联合训练共享策展器,利用检索感知和行为感知的时间归因,将优化聚焦于具有可观察下游影响的策展动作。在ALFWorld、ScienceWorld和WebShop上,执行器范围从Qwen3-8B/32B和GPT-OSS-120B到未见过的Kimi K2.6、DeepSeek V4 Flash和Gemini 3.5 Flash,EASE在无需逐执行器微调的情况下,优于强技能和记忆基线。EASE还保持技能数量减少34.5%至41.0%,技能检索提升36.3%至38.7%,测量编辑效用提升51.8%至60.0%,并将部署时推理令牌减少9.1%至14.5%。这些结果确立了行为自适应技能策展作为构建自进化智能体的有效原则。

英文摘要

Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explicitly modeling downstream executor behavior. We show that this can cause systematic cross-executor degradation: curators trained with different executors perform best when paired with their own training executor, indicating that effective skill curation is executor-dependent. We formulate behavior-adaptive skill curation and introduce EASE, a framework that learns a single curator that adapts its decisions to different executor behaviors. EASE maintains an online behavioral profile of recent execution patterns and conditions the curator on this profile, the current trajectory, and retrieved skills to add, modify, or remove skills from an evolving repository. We train the shared curator jointly across multiple frozen executors with reinforcement learning, using retrieval-aware and behavior-aware temporal attribution to focus optimization on curation actions with observable downstream influence. Across ALFWorld, ScienceWorld, and WebShop, with executors ranging from Qwen3-8B/32B and GPT-OSS-120B to unseen Kimi K2.6, DeepSeek V4 Flash, and Gemini 3.5 Flash, EASE outperforms strong skill- and memory-based baselines without per-executor finetuning. EASE also maintains 34.5--41.0% fewer skills, improves skill retrieval by 36.3--38.7% and measured edit utility by 51.8--60.0%, and reduces deployment-time inference tokens by 9.1--14.5%. These results establish behavior-adaptive skill curation as an effective principle for building self-evolving agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑