TGT:用于局部可控视频生成的文本锚定轨迹
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
浏览论文内容
中文总结 AI 辅助
本文提出TGT框架,通过将点轨迹与局部文本描述配对并采用位置感知交叉注意力和双CFG方案,实现视频生成中外观和运动的局部可控,实验证明其优于现有方法。
中文摘要 AI 辅助
文本到视频生成在视觉保真度方面取得了快速进展,而标准方法在控制生成场景的主体构成方面能力仍然有限。先前的工作表明,添加局部文本控制信号(如边界框或分割掩码)可以有所帮助。然而,这些方法在复杂场景中表现不佳,在多对象设置中性能下降,提供的精度有限,并且随着可控对象数量的增加,单个轨迹与视觉实体之间缺乏清晰的对应关系。我们引入了文本锚定轨迹(TGT),这是一个将视频生成条件于与局部文本描述配对的轨迹上的框架。我们提出了位置感知交叉注意力(LACA)来整合这些信号,并采用双CFG方案分别调制局部和全局文本引导。此外,我们开发了一个数据处理流水线,生成带有跟踪实体局部描述的轨迹,并标注了两百万个高质量视频片段来训练TGT。这些组件共同使TGT能够使用点轨迹作为直观的运动句柄,将每个轨迹与文本配对以控制外观和运动。大量实验表明,与先前方法相比,TGT实现了更高的视觉质量、更准确的文本对齐和改进的运动可控性。网站:https://textgroundedtraj.github.io。
英文摘要
Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior work shows that adding localized text control signals, such as bounding boxes or segmentation masks, can help. However, these methods struggle in complex scenarios and degrade in multi-object settings, offering limited precision and lacking a clear correspondence between individual trajectories and visual entities as the number of controllable objects increases. We introduce Text-Grounded Trajectories (TGT), a framework that conditions video generation on trajectories paired with localized text descriptions. We propose Location-Aware Cross-Attention (LACA) to integrate these signals and adopt a dual-CFG scheme to separately modulate local and global text guidance. In addition, we develop a data processing pipeline that produces trajectories with localized descriptions of tracked entities, and we annotate two million high quality video clips to train TGT. Together, these components enable TGT to use point trajectories as intuitive motion handles, pairing each trajectory with text to control both appearance and motion. Extensive experiments show that TGT achieves higher visual quality, more accurate text alignment, and improved motion controllability compared with prior approaches. Website: https://textgroundedtraj.github.io.
发表机构
- Johns Hopkins University(约翰霍普金斯大学)
- Bytedance, Intelligent Creation(字节跳动,智能创作)
机构由 AI 辅助整理,请以论文原文为准。