发表机构
State Key Laboratory of CAD&CG, Zhejiang University; HKUST(GZ); LIGHTSPEED(浙江大学CAD&CG国家重点实验室; 香港科技大学(广州); LIGHTSPEED)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出GEAR框架,利用几何作为令牌级地址路由注意力到视觉记忆,避免全局融合误差,实现长时程相机控制视频生成的高质量与一致性。
AI 中文摘要
长时程相机控制视频生成需要从不断增长的视觉历史中恢复先前观察到的内容。现有方法要么隐式搜索历史上下文,要么将其重建为持久的三维记忆,但面临低效的内存访问或累积的几何误差。我们的关键洞见是,几何无需解释场景——它只需确定视觉记忆应从何处读取,而注意力决定应恢复什么内容。基于此洞见,我们提出了GEAR,一种几何启用的注意力路由框架,利用几何作为视觉记忆的显式令牌级地址。GEAR不将历史观测融合到持久的全局三维表示中,而是将其保留为帧潜变量,并仅使用每帧几何来建立与目标视图的令牌级对应关系,从而避免全局融合带来的持久误差累积。在这些对应关系的引导下,几何对应注意力(GCA)在去噪过程中选择性地将几何匹配的历史特征注入到噪声目标块中。我们进一步引入了一种不可见八叉树来累积可见性证据,并拒绝几何上合理但被遮挡的对应关系。大量实验表明,GEAR在视觉质量、精确相机控制和重访一致性方面达到了最先进水平,能够沿具有挑战性的轨迹生成分钟级视频。
英文摘要
Long-horizon camera-controlled video generation requires recovering previously observed content from an ever-growing visual history. Existing approaches either search historical context implicitly or reconstruct it into persistent 3D memory, facing inefficient memory access or accumulated geometric errors. Our key insight is that geometry need not explain the scene--it only needs to determine where visual memory should be read from, while attention decides what should be recovered. Based on this insight, we introduce GEAR, a Geometry-Enabled Attention Routing framework that uses geometry as an explicit token-level address for visual memory. Rather than fusing historical observations into a persistent global 3D representation, GEAR retains them as frame latents and uses per-frame geometry only to establish token-level correspondences with target views, thereby avoiding persistent error accumulation from global fusion. Guided by these correspondences, Geometric Correspondence Attention (GCA) selectively injects geometrically matched historical features into noisy target patches during denoising. We further introduce an Invisible Octree to accumulate visibility evidence and reject geometrically plausible but occluded correspondences. Extensive experiments demonstrate that GEAR achieves state-of-the-art visual quality, precise camera control, and revisit consistency, enabling minute-long video generation along challenging trajectories.
CommentsProject Page: https://zju3dv.github.io/geometry-as-address/