发表机构
University of Cambridge; Shanghai Jiao Tong University; University of Twente; University of Warwick(剑桥大学; 上海交通大学; 特温特大学; 华威大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有多帧场景流估计方法的计算开销大、长时程预测性能差的问题,提出MESSENGER模型,通过内存缓冲区和不确定性感知重加权模块,在nuScenes和Argoverse 2上实现了最先进的长时程外推性能。
AI 中文摘要
场景流可捕捉动态场景中的低级三维运动位移。早期的成对估计器依赖瞬时两帧运动,缺乏长期时间相关性,且在未来预测中存在外推能力差的问题。尽管近期一些方法尝试以序列到序列的方式探索多帧场景流估计,但随着输入帧数增加,它们通常存在计算开销大的问题,且因运动传播无效而导致长时程预测性能下降。为解决这些问题,我们提出了一种名为MESSENGER的新型内存增强型序列场景流管道。为充分挖掘连续序列中自然存在的长期时间依赖关系,我们设计了一个内存缓冲区,用于显式存储多个历史流估计值和隐状态。对于每个输入帧,我们将时间存储的流和状态进行关联并检索,以以下一帧预测的方式预测当前初始化流。此外,我们开发了一个不确定性感知重加权模块,用于过滤不可靠的检索结果并缓解累积误差。在nuScenes和Argoverse 2上进行的大量实验表明,我们的MESSENGER达到了最先进的性能,在长时程未来外推任务中,nuScenes上的EPE3D降低了71.6%,Argoverse 2上的EPE3D降低了67.7%。这种优势可归因于我们设计的自回归预测范式,该范式自然地迫使网络基于历史观测逐步学习下一帧的分布。代码将在此httpsURL发布。
英文摘要
Scene flow can capture low-level 3D motion displacements in dynamic scenarios. Early pairwise estimators relying on instantaneous two-frame motion lack long-term temporal correlation and also struggle with poor extrapolation ability in future prediction. Although some recent methods attempt to explore multi-frame scene flow estimation in a sequence-to-sequence manner, they typically suffer from heavy computational overhead with increasing input frames and long-horizon prediction degradation due to ineffective motion propagation. To address these problems, we propose a novel memory-enhanced sequential scene flow pipeline, called MESSENGER. To sufficiently mine long-term temporal dependencies naturally within consecutive sequences, a memory buffer is designed by explicitly storing multiple history flow estimates and latent states. For each input frame, the temporally stored flows and states are correlated and retrieved to predict the current initialized flow in a next-frame forecasting manner. Furthermore, we develop an uncertainty-aware reweighting module to filter unreliable retrievals and mitigate accumulated errors. Extensive experiments on nuScenes and Argoverse 2 demonstrate state-of-the-art performance of our MESSENGER, reducing EPE3D by 71.6% on nuScenes and 67.7% on Argoverse 2 in long-horizon future extrapolation. This superiority can be attributed to our designed autoregressive forecasting paradigm, which naturally forces the network to progressively learn the next-frame distribution based on history observations. Code will be released at https://github.com/liujiuming123/Messenger.
CommentsAccepted by NeurIPS 2026. Code will be released at: https://github.com/liujiuming123/Messenger