arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在流中做梦:面向自进化统一多模态模型的生成式接地反馈

Dreaming in Flow: Generative Grounding Feedback for Self-Evolving Unified Multimodal Models

Ke Hao, Yuanzhi Liang, Tingxi Chen, Rui Li, Haibin Huang, Chi Zhang, Yun Gu, Xuelong Li

arXiv 2609.08282首次发表:更新:

发表机构

Shanghai Jiao Tong University; Institute of Artificial Intelligence, China Telecom (TeleAI); University of Science and Technology of China(上海交通大学; 中国电信人工智能研究院(TeleAI); 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出生成式接地反馈(GGF)框架,通过流级反馈和梦境重放接地,利用模型自身视觉经验实现自进化,无需配对监督,提升统一多模态模型的生成与理解性能。

AI 中文摘要

统一多模态模型在单一网络内整合了视觉理解与生成能力,然而这两种能力通常被作为独立任务进行优化。我们提出了生成式接地反馈(Generative Grounding Feedback, GGF),一种自进化的训练后框架,仅使用文本提示和模型自身的视觉经验。给定一个提示,模型首先生成一个视觉“梦境”。流级反馈在相同的噪声潜状态上比较文本条件、图像条件和修复条件的预测,将图像接地的生成方向迁移到提示条件上。梦境重放接地通过描述和再想象来重放这一梦境,训练声明级证据在重放过程中保持一致,同时分离无关的视觉经验。通过联合优化,这两个方向使生成为理解提供视觉接地,而理解则细化后续的生成,无需配对的图像-文本监督。在具有不同理解-生成集成设计的统一模型上的实验表明,文本到图像生成性能持续提升,同时视觉理解能力也有适度提高。

英文摘要

Unified multimodal models integrate visual understanding and generation within a single network, yet the two capabilities are commonly optimized as separate tasks. We introduce Generative Grounding Feedback(GGF), a self-evolving post-training framework that uses only text prompts and the model's own visual experience. Given a prompt, the model first generates a visual ``dream.'' Flow-level feedback compares text-, image-, and repair-conditioned predictions at the same noisy latent state, transferring image-grounded generation directions to the prompt condition. Dream replay grounding replays this dream through captioning and re-imagination, training claim-level evidence to remain consistent across the replay while separating unrelated visual experiences. Jointly optimized, these two directions let generation provide visual grounding for understanding and understanding refine subsequent generation without paired image--text supervision. Experiments across unified models with different understanding--generation integration designs show consistent improvements in text-to-image generation together with modest gains in visual understanding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑