基于多模态大语言模型(MLLM)的文本到视频生成语义修正
MLLM-Guided Semantic Correction for Text-to-Video Generation
浏览论文内容
中文总结 AI 辅助
该研究提出一种无需训练的MLLM引导文本到视频生成语义修正框架,通过两个关键模块在生成中修正语义偏差,提升了生成内容的语义对齐度等性能,经多基准实验验证有效。
中文摘要 AI 辅助
扩散模型与Transformer架构的近期进展推动了文本到视频生成技术的显著进步,但这类模型常存在语义错误,如对象缺失、属性错误或动作不匹配。尽管部分语义修正方法在采样前优化或采样后细化,但在视频生成过程中检测并修正语义偏差的研究仍较欠缺。本文提出一种无需训练、可解释的生成中修正框架,将多模态大语言模型(MLLM)反馈直接融入扩散采样循环,通过在视频合成时注入语义评估信号实现扩散轨迹修正,让模型通过持续自我反思优化生成内容。该框架包含两个关键模块:语义评估监督器,生成中间预览帧用于语义评估与偏差诊断;语义修正助手,在推理时通过可控潜在轨迹干预修正语义漂移。本文方法无需修改模型参数,即可提升语义对齐度、视觉保真度与时间一致性,通过多个基准的大量实验验证了其有效性。
英文摘要
Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic errors such as missing objects, incorrect attributes, or mismatched actions. Although some semantic correction methods perform optimization before sampling or refinement after sampling, how to detect and correct semantic deviations during the video generation process remains underexplored. In this paper, we introduce a training-free, interpretable mid-generation correction framework that integrates multimodal large language model (MLLM) feedback directly into the diffusion sampling loop. Our framework achieves diffusion trajectory correction by injecting semantic evaluation signals during video synthesis, enabling the model to optimize the generated content through continuous self-reflection. We propose two key modules: a Semantic Assessment Supervisor that generates intermediate preview frames for semantic evaluations and deviation diagnostics, and a Semantic Modification Assistant that corrects semantic drift during inference via a controllable latent trajectory intervention. Our method improves semantic alignment, visual fidelity, and temporal consistency without modifying model parameters. We validate the effectiveness of our approach through extensive experiments across multiple benchmarks.
发表机构
- Zhejiang University(浙江大学)
- Huawei Cloud Computing Technology Co., Ltd.(华为云计算技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。