用于动态新视角合成的在线神经时空记忆
Online Neural Space Time Memory for Dynamic Novel View Synthesis
浏览论文内容
中文总结 AI 辅助
研究多视图流视频在线新视角合成问题,提出解耦记忆更新与应用频率的方法,通过跨视图注意力管理变形,引入辅助记忆损失和记忆缓存策略,实现实时、领先性能及微小尺度在线记忆。
中文摘要 AI 辅助
从多视图流视频进行在线新视角合成面临一个基本权衡:在严格的实时约束下运行时,既要维护持久的长时记忆以重建暂时遮挡的区域。虽然测试时训练(TTT)提供了强大的记忆机制,但标准模型要求在每一帧基于梯度更新记忆以适应动态场景中变化的运动。大量记忆更新的计算成本排除了实时应用,且可能导致长上下文的不稳定。鉴于记忆更新比记忆应用要求更高且视频内容大多冗余,我们提出解耦这两个过程的频率。我们的方法在逐帧应用记忆时进行周期性记忆更新,使用跨视图注意力管理先前记忆状态与当前帧之间的变形。为锁定历史上下文,我们引入两个关键机制:辅助记忆损失强制场景的持久内化,以及记忆缓存策略规范活动权重以防灾难性漂移。我们的方法在具有动态人体运动的场景以及微小尺度的在线记忆方面展示了实时的、领先的性能。
英文摘要
Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models mandate gradient-based memory updates at every frame to adapt to the changing motion in dynamic scenes. The computational cost of heavy memory updates precludes real-time application and can lead to instability over long contexts. Given that memory updates are more demanding than memory application and video content is largely redundant, we propose to decouple the frequencies of these two processes. Our approach performs periodic memory updates while applying the memory on a per-frame basis, using cross-view attention to manage deformations between the prior memory state and the current frame. To lock in the historical context, we introduce two critical mechanisms: an auxiliary Memory Loss that forces persistent internalization of the scene, and a Memory Caching strategy that regularizes active weights against catastrophic drift. Our method demonstrates state-of-the-art minute-scale memory persistence in online dynamic human scenes at amortized real-time speed.
发表机构
- University of Washington(华盛顿大学)
- Google(谷歌)
机构由 AI 辅助整理,请以论文原文为准。