发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于状态空间模型(Mamba-3)的三维点跟踪方法,结合光流与单目深度网络,在单GPU预算下实现绝对度量精度领先(TAPVid-3D平均Jaccard 0.256)。
AI 中文摘要
在度量三维空间(以绝对米为单位,而非未知尺度)中跟踪动态场景的任意点,是三维和四维重建、机器人导航及自动驾驶的基础,这些场景中的决策以米而非像素为单位。我们的目标是构建一个三维点跟踪器,在绝对度量意义上保持准确,并能在单块商用GPU上运行,无需位姿信息,仅依赖单目输入。我们的方法基于一个观察:一旦点的二维图像轨迹确定,决定其度量精度的关键量是沿像素射线的深度。因此,我们不采用端到端的跟踪学习,而是组合两个冻结的前端——用于二维对应的稠密光流和用于第三维度的单目度量深度网络——仅学习它们无法提供的残差:即由紧凑状态空间模型(Mamba-3)基于外观特征(DINOv3)精化的深度。采用状态空间模型而非最强三维跟踪器所用的Transformer,使得单GPU预算可行:它通过固定大小的循环状态汇总轨迹,其内存成本在帧数上恒定,而注意力机制需要随帧数线性增长的键值缓存。在TAPVid-3D minival基准上,我们最佳配置在相同条件下评估的方法中取得了最高的绝对度量精度(平均度量Jaccard指数,0.256),超越了强前馈跟踪器;同时,一项伴随分析(使用每个竞争者的评估器复现)解释了为何若干已发表的跟踪器在此预算下会损失大部分精度。
英文摘要
Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget. Our method rests on one observation: once a point's 2D image trajectory is fixed, the quantity that governs its metric accuracy is the depth along its pixel ray. Rather than learning tracking end-to-end, we therefore compose two frozen front-ends - dense optical flow for 2D correspondence and a monocular metric-depth network for the third dimension - and learn only the residual they cannot supply: that depth, refined by a compact state space model (Mamba-3) conditioned on appearance features (DINOv3). A state space model rather than the transformers the strongest 3D trackers adopt is what makes a single-GPU budget attainable: it summarises a track in a fixed-size recurrent state whose memory cost is constant in the number of frames, whereas attention requires a key-value cache that grows linearly with them. On the TAPVid-3D minival benchmark our best configuration attains the highest absolute metric accuracy among methods evaluated under identical conditions (mean metric Average Jaccard, 0.256), exceeding strong feed-forward trackers, while a companion analysis, reproduced with each competitor's own evaluator, explains why several published trackers lose most of their accuracy under this budget.
Comments20 pages, 11 figures, 9 tables. Submitted to Computer Vision and Image Understanding