arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07250cs.AI

通过在线策略上下文蒸馏将智能体经验内化到扩散模型权重中

Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation

Wenxuan Wang, Zekai Liu, Weinan Zhang, Yu Cheng, Yang Yang

首次发表
浏览论文内容

中文总结 AI 辅助

提出D-OPCD方法,将智能体改进提示词作为特权上下文蒸馏进扩散模型权重,使模型在无框架时保留部分增益,在四个基准上平均得分从60.52升至65.09,并支持框架与模型持续共同进化。

中文摘要 AI 辅助

将图像生成模型包装在智能体框架中能有效提升文生图任务性能:该框架可以利用记忆、技能、工作流编排、结果验证和迭代细化来持续构建和修改提示词,从而生成更优质的图像。然而,这些增益仍然外在于扩散模型,仅在完整框架运行时才能实现。我们提出扩散模型在线策略上下文蒸馏(D-OPCD),该方法将智能体改进后的提示词视为特权上下文,并将智能体框架中编码的知识蒸馏到扩散模型的权重中,使得模型在仅基于原始查询的条件下也能保留框架的部分收益。使用配备我们提出的自动技能进化器(ASE)的文生图智能体,我们证明D-OPCD能将框架能力内化到生成器的权重中,在四个基准上将平均直接生成分数从60.52提升至65.09。随着这些知识被吸收进权重,框架可以舍弃饱和的技能并继续进化:在更新后的生成器上进行第二轮ASE,相比无技能框架额外提升了1.83分,这指向了框架与模型通过持续共同进化不断相互提升的文生图系统。

英文摘要

Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however, remain external to the diffusion model and are realized only while the full harness runs. We propose Diffusion On-Policy Context Distillation (D-OPCD), which treats the agent-improved prompt as privileged context and distills the knowledge encoded in the agent harness into the weights of the diffusion model, so that the model retains part of the harness's benefit when conditioned on the original query alone. Using a Text-to-Image agent equipped with our proposed Auto Skill Evolver (ASE), we show that D-OPCD can internalize harness capabilities into the generator's weights, raising the average direct-generation score from 60.52 to 65.09 across four benchmarks. With this knowledge absorbed into the weights, the harness can shed its saturated skills and resume evolving: a second ASE round on the updated generator improves on a skill-free harness by additional 1.83 points, pointing toward text-to-image systems in which harness and model keep improving each other through continual co-evolution.

发表机构

  • Harbin Institute of Technology(哈尔滨工业大学)
  • Shanghai AI Laboratory(上海人工智能实验室)
  • Shandong University(山东大学)
  • Nanyang Technological University(南洋理工大学)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑