arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过自监督语义扩散将技能像参数一样训练

Skill Training with Corruption and Reconstruction Loop

Mo Li, Zixin Yin, Qihao Wu, Ting Cao, Yunxin Liu, Heung-Yeung Shum

arXiv 2607.27557首次发表:更新:

发表机构

Tsinghua University; Shanghai AI Laboratory; The Hong Kong University of Science and Technology(清华大学; 上海人工智能实验室; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出受扩散模型启发的无监督自进化智能体框架,通过自监督信号更新外部文本技能库,提升了智能体在短剧剧本创作领域的生成能力。

AI 中文摘要

尽管大型语言模型(LLMs)展现出出色的通用指令遵循能力,但在创意剧本创作等高度专业化、开放性的领域,它们往往不及人类专家。现有方法通常采用后训练方式,然而监督微调与强化学习均需权重访问权限,而闭源前沿模型并不提供该权限,且二者均需大量计算资源。此外,所学内容与单个检查点绑定,人类无法对其进行检查。近期的智能体持续学习进展试图通过积累外部文本技能来弥合这一差距,但这些方法严重依赖成本高昂的人类专家标注或不可靠的LLM作为评判者的反馈来进行反思。为克服这一瓶颈,我们提出一种受扩散模型的损坏-重建范式启发的新型无监督自进化智能体框架。该框架不依赖显式外部评分,而是利用现有高质量人类作品构建自监督信号,训练遵循神经网络训练的常见循环:前向传播、损失计算与反向传播,其中损失来自将智能体的重建结果与人类原作进行对比。更新的对象并非模型权重,而是外部文本技能库。我们在具有挑战性的短剧剧本创作任务上评估该框架,实验结果表明,我们的方法使智能体能够自主提取并内化高度可泛化的技能,显著提升其特定领域的生成能力。此外,这种自对比反思范式为智能体提供了可扩展的途径,使其无需外部监督即可自学生成复杂、高质量的人类作品。

英文摘要

Large Language Models (LLMs) often struggle in highly specialized domains. Rather than parameter-level adaptation of LLMs which is costly and difficult to interpret, external skills (often defined as text files) have been recently proposed to augment LLMs for specialized domains. However, such skills rely on costly active human annotations or passive summarization of high-quality examples. In this paper, we propose a self-supervised approach for agent self-evolution that learns domain-specific skills directly from existing high-quality human artifacts, without additional human annotations or external rewards. Inspired by diffusion models, our approach follows a forward-loss-backward process to reconstruct human artifacts by iteratively learning the agent's external skill library rather than updating its model parameters. Experiments on short-drama screenwriting demonstrate that our approach enables agents to autonomously extract generalizable writing skills from human-authored scripts and substantially improve domain-specific generation quality. Our approach provides a scalable paradigm for agents to continuously learn many kinds of complex skills from existing high-quality human artifacts. Code and project page: https://github.com/skilltraining-project/skill-training and https://skilltraining-project.github.io

CommentsUnder review. Code: https://github.com/skilltraining-project/skill-training

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑