arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15667cs.CV

PE-Field 4D:作为画布的视频生成模型

PE-Field 4D: Video Generation Models as Canvas

发表机构德克萨斯大学奥斯汀分校 · 皮克社交科技
查看机构详情
  • University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Pixocial Technology(皮克社交科技)

机构由 AI 辅助整理,请以论文原文为准。

Yunpeng Bai, Haoxiang Li, Qixing Huang

首次发表
浏览论文内容

中文总结 AI 辅助

研究视频生成中扩散变换器控制场景几何形状的挑战,通过重新审视位置编码作用,引入几何感知交叉注意力机制及相关编码方案,构建可控视频生成框架,提升视点相关编辑任务的空间可控性且保留模型生成先验。

中文摘要 AI 辅助

扩散变换器最近在视频生成中取得了强大性能,但在视点变化和相机运动下控制场景几何形状仍具有挑战性。本文重新审视了视频扩散变换器中位置编码的作用,表明其为几何感知控制提供了有用的空间偏差。具体而言,若参考令牌根据其在目标视图中的投影位置进行编码,去噪模型会从输入视频的位置对齐区域检索内容。在此基础上,引入了几何感知交叉注意力机制,使目标视频潜在令牌能关注从参考图像或帧派生的结构化上下文令牌。为在参考内容与目标相机轨迹间建立对应关系,为上下文令牌配备了投影位置编码方案,结合目标视图二维重投影与深度感知消歧。同时保留生成视频潜在的原始时空位置编码,使几何引导得以注入,同时保持与视频模型原生潜在结构的一致性。由此产生的框架为可控视频生成提供了简单有效的方法,在视点相关编辑任务中提高了空间可控性,包括相机重轨迹、新视图视频合成和几何感知视频编辑,同时保留了底层视频扩散模型的生成先验。代码可在该https网址获取。

英文摘要

Diffusion Transformers have recently achieved strong performance in video generation, yet controlling scene geometry under viewpoint changes and camera motion remains challenging. In this work, we revisit the role of positional encoding in video diffusion transformers and show that it provides a useful spatial bias for geometry-aware control. Specifically, if reference tokens are encoded according to their projected locations in the target view, the denoising model is encouraged to retrieve content from position aligned regions of the input video. Building on this observation, we introduce a geometry-aware cross-attention mechanism that enables target video latent tokens to attend to structured context tokens derived from reference images or frames. To establish correspondence between the reference content and the target camera trajectory, we equip the context tokens with a projected positional encoding scheme that combines target-view 2D reprojection with depth-aware disambiguation. At the same time, we preserve the original spatiotemporal positional encoding of the generated video latent, allowing geometric guidance to be injected while maintaining consistency with the video model's native latent structure. The resulting framework provides a simple and effective approach for controllable video generation. It improves spatial controllability in viewpoint-dependent editing tasks, including camera re-trajectory, novel-view video synthesis, and geometry-aware video editing, while preserving the generative prior of the underlying video diffusion model. The code is available at: https://github.com/MTLab/PE-Field.

↑