LoGo:用于一致长时程视频生成的局部-全局奖励
LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation
浏览论文内容
中文总结 AI 辅助
针对相机控制视频生成中的三维不一致问题,提出融合局部与全局奖励的LoGo方法,通过细粒度信用分配提升长时程一致性,在DL3DV和TrajectoryBench基准上优于现有方法。
中文摘要 AI 辅助
相机控制的视频模型正迅速向长生成时程和复杂相机控制发展。一个关键的失败模式是三维不一致性:随着相机移动,物体失去持久性,场景结构发生偏移。现有的后训练技术将单一标量奖励分配给整个生成过程,难以在长时程中纠正这些不一致性。我们提出LoGo,它将全局奖励与空间局部奖励融合用于相机控制的视频模型。局部奖励提供细粒度的信用分配,显著提升三维一致性,而全局奖励则保持相机跟随和视频质量。在三个基础模型上,LoGo在DL3DV和TrajectoryBench(一个针对当前评估缺乏的长时程、复杂相机控制生成的新基准)上展现出明显优势。LoGo有效减少了局部物体偏移、伪影和全局场景变化,突显了后训练视频模型中信用分配的重要性。项目网站:此https URL
英文摘要
Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video models. The local reward provides fine-grained credit assignment, which substantially improves 3D consistency, while the global reward preserves camera following and video quality. Across three base models, LoGo shows a clear advantage on DL3DV and TrajectoryBench, a new benchmark for long-horizon, complex-camera-control generation that current evaluations lack. LoGo effectively reduces local object shifts, artifacts, and global scene changes, illustrating the importance of credit assignment in post-training video models. Project website: https://ziqi-ma.github.io/logo-website/
发表机构
- California Institute of Technology(加州理工学院)
- World Labs
机构由 AI 辅助整理,请以论文原文为准。