arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillGLoW:面向长周期任务流的自改进智能体的过程族技能巩固

SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

Ao Yan, Xin Zhang, Jiawei Du, Joey Tianyi Zhou

arXiv 2609.02217首次发表:更新:

发表机构

National University of Singapore; Institute of Advanced Intelligence and Computing (IAIC)(新加坡国立大学; 先进智能与计算研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出SkillGLoW,通过聚合任务执行生成的局部技能为过程族并压缩为全局先验,在四类基准和三个模型上提升了智能体的长周期任务性能,且实现了紧凑的技能库与任务间的过程迁移。

AI 中文摘要

大型语言模型(LLM)智能体正通过编写和复用文本技能实现日益显著的自改进,这些技能要么以单个全局文档形式存储,要么以每个任务条目的扁平池形式存储,不过多数相关证据来自任务结构相似的领域。在每个任务需要不同解决方案的长周期工作负载场景下,这两种形式会以相反方式失效:全局文档会退化为通用规范,而扁平池会不断膨胀,且其条目始终绑定到编写它们的实例。本文认为缺失的复用单元是一组相关任务共享的求解过程,并围绕该单元构建了SkillGLoW(Global-Local Weave,全局-局部编织):任务从自身执行中生成的局部技能被聚合为过程族,并被压缩为去实例化的全局先验,它们所包含的实例细节会针对每个任务重新生成,而非存储;提交门仅在实际执行显示先验不会降低已部署库性能时才允许该先验加入。在四个基准(数学推理、终端自动化、软件修复、具身控制)和三个模型上,这些先验相比无技能基线的平均提升为17.2个百分点(针对困难任务),在12次持续改进运行中均实现正向提升,结合局部再生后提升达18.0个百分点,同时该库每个过程族仅保留一个先验,比按任务划分的扁平池紧凑3.6倍。在相同协议下,GLoW在21个单元中的15个上领先已发表的单文档优化器。未修改的该库将未见过的ALFWorld任务的成功率从73.9%提升至83.9%,证明可迁移的是过程而非任务记忆。

英文摘要

LLM agents increasingly self-improve by writing and reusing textual skills, kept either as one global document or as a flat pool of per-task entries, though most of the evidence comes from domains with structurally similar tasks. On long-horizon workloads where each task demands a different solution, the two forms fail in opposite ways: the document collapses into generic discipline, while the pool inflates and its entries stay bound to the instance that wrote them. We argue the missing unit of reuse is the solving procedure shared by a cluster of related tasks, and build SkillGLoW (Global-Local Weave) around it: the local skills a task writes from its own execution are aggregated into procedural families and compressed into de-instantiated global priors, while the instance detail they hold is regenerated per task rather than stored; a commit gate admits a prior only when real execution shows it does not degrade the deployed library. Across four benchmarks (mathematical reasoning, terminal automation, software repair, and embodied control) and three models, the priors gain 17.2 points (hard) over the no-skill baseline on average, with positive gains in all 12 continual-improvement runs, and 18.0 with local regeneration, while the library holds one prior per procedural family, 3.6x more compact than the per-task pool. Under the same protocol GLoW leads a published single-document optimizer on 15 of 21 cells. Unmodified, the library lifts success on unseen ALFWorld tasks from 73.9% to 83.9%, evidence that what transfers is procedure rather than task memory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑