交错多模态生成的自我修正优化
Self-correction Optimization for Interleaved Multimodal Generation
查看机构详情
- Shanghai Jiao Tong University(上海交通大学)
- SenseTime Research(商汤科技研究院)
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出无需训练的自我修正优化方法,通过新事件与状态保持约束,提升交错多模态生成的时间一致性与视觉主体保持,并适用于视频生成。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)在视觉理解和生成方面取得了显著进展。然而,生成交错的图像-文本内容仍然具有挑战性,因为这需要紧密集成的多模态理解和生成能力。尽管现有的MLLMs提供了有前景的解决方案,但大多数依赖于对增强数据的额外训练,这计算成本高昂,并且在保持视觉主体、时间一致性和物理合理性方面仍然有限。在这项工作中,我们提出了自我修正优化(SCO),一种有效的无需训练的方法,用于一致的交错生成。SCO将无分类器引导更新视为参考,并在两个互补约束下执行最小自我修正,包括新事件和状态保持约束。具体来说,新事件约束促进了图像-文本序列间的时间一致性,而状态保持约束在后续生成步骤中维持了视觉主体的连贯性。在具有挑战性的交错多模态生成基准上的实验表明,在时间连贯性和视觉主体保持方面有显著改进。此外,SCO可以扩展到视频生成,并改进对物理基础过程的建模,包括机器人操作和长时间手工制作。
英文摘要
Multimodal large language models (MLLMs) have made significant progress in visual understanding and generation. However, generating interleaved image--text content remains challenging, as it requires tightly integrated multimodal understanding and generation capabilities. Although existing MLLMs provide promising solutions, most rely on additional training with augmented data, which is computationally expensive and remains limited in preserving visual subjects, temporal consistency, and physical plausibility. In this work, we propose self-correction optimization (SCO), an effective training-free method for consistent interleaved generation. SCO treats the classifier-free guidance update as a reference and performs minimal self-correction under two complementary constraints, including new-event and state-preserving constraints. Specifically, the new-event constraint promotes temporal consistency across image--text sequences, while the state-preserving constraint maintains the coherence of visual subjects throughout subsequent generation steps. Experiments on challenging interleaved multimodal generation benchmarks demonstrate significant improvements in temporal coherence and visual-subject preservation. Furthermore, SCO can be extended to video generation and improves the modeling of physically grounded processes, including robot manipulation and long-horizon handcrafting.