arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Rep2Skill:面向LLM智能体的表示引导技能自进化

Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

Kaixing Zhang, Changming Li, Yingdong Shi, Zheng Zhang, Kaitao Song, Wenjie Shi, Jingang Wang, Kan Ren

arXiv 2609.39149首次发表:更新:

发表机构

ShanghaiTech University; Meituan(上海科技大学; 美团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Rep2Skill通过建模智能体内部表示轨迹定位偏离成功执行的回合,生成文本反馈以修订技能,在自进化设置中优于纯文本方法,推动智能体自我改进超越文本反思。

AI 中文摘要

文本技能使基于大型语言模型(LLM)的智能体能够在无需更新模型参数的情况下积累可复用的程序性知识。然而,现有的技能进化在很大程度上仍局限于文本空间:优化器必须诊断成功与失败的模式,并且仅从冗长的执行轨迹和稀疏的任务结果中修订技能。这种纯文本范式将智能体的内部表示(其中包含其不断演化的执行状态的丰富记录)排除在技能优化循环之外。我们提出一个问题:智能体能否通过反思自身的内部表示来改进其外部文本技能?我们引入了Rep2Skill,一个用于智能体技能自进化的表示引导框架。具体而言,在收集到的智能体轨迹上,Rep2Skill对其内部模型表示轨迹进行建模,以定位偏离成功执行动态的回合,并进一步将这些信号与执行上下文一起解释为可操作的文本反馈,用于有针对性的技能修订。在两个智能体环境上使用两个开源LLM进行的实验表明,在自进化设置中(同一LLM同时充当执行器和优化器,且没有更强的外部模型),Rep2Skill始终优于纯文本方法。这为将智能体自我改进超越纯文本反思开辟了一个有前景的方向。

英文摘要

Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must diagnose success and failure patterns, and revise skills solely from long execution trajectories and sparse task outcomes. This text-only paradigm leaves the agent's internal representations, which contain rich records of its evolving execution state, outside the skill optimization loop. We ask whether an agent can improve its external textual skills by reflecting on its own internal representations. We introduce Rep2Skill, a representation-guided framework for self-evolution on agent skills. Specifically, upon the collected agent rollouts, Rep2Skill models their internal model representation trajectories to localize turns that deviate from successful execution dynamics, and it further interprets these signals alongside the execution contexts as actionable textual feedback for targeted skill revision. Experiments on two agent environments with two open-source LLMs show that Rep2Skill consistently outperforms text-only approaches in the self-evolution setting, where the same LLM serves as both executor and optimizer without a stronger external model. This establishes a promising direction moving agent self-improvement beyond text-only reflection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑