发表机构
Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RegVGGT提出无需训练的token调控方法,通过每帧仅保留1%显著token更新内存,在消费级GPU上实现长视频流的高效三维重建,性能大幅超越现有基线。
AI 中文摘要
从长视频流输入进行三维重建,对前馈重建模型(FFRMs)构成了一个两难问题:在有限的GPU内存下,无法保留整个流的推理上下文。先前的研究试图通过权衡推理上下文的完整性与GPU内存使用来解决这一问题,但这些方法要么遭受内存快速膨胀,要么因人为限制内存而导致上下文完整性下降。基于我们的关键观察——即一个token的初始显著性可靠地决定了其在整条流中的长期重要性,我们提出了RegVGGT,一种无需训练的token调控方法,该方法对传入帧的token进行激进调控。通过每帧仅允许最多1%的token更新上下文内存,我们的方法显著抑制了随流增长的内存膨胀。结合与FlashAttention兼容的token显著性估计方案,RegVGGT能够在消费级GPU上处理数千帧,且对重建质量的影响可忽略不计。我们的实验表明,RegVGGT在多种FFRM预测任务的长时程基准上取得了最先进的性能,大幅超越了先前基于FFRM的流重建基线。
英文摘要
3D reconstruction from a lengthy video stream input poses a dilemma for feed-forward reconstruction models (FFRMs), that a whole-stream inference context cannot be retained under limited GPU memory.Recent studies seek to resolve this problem via a trade-off between the integrity of inference context and GPU memory usage, which either suffer from a rapid memory inflation or degraded context integrity due to artificially capping memory usage.Driven by our key observation that the initial saliency of a token reliably dictates its long-term importance across the stream, we propose RegVGGT, a training-free token regulation method which aggressively regulates the tokens of incoming frames.By admitting at most 1% of tokens per frame to update the context memory, our method dramatically suppresses memory inflation as the stream progresses.Equipped with a FlashAttention-compatible token saliency estimation scheme, RegVGGT is capable of processing thousands of frames on a consumer-grade GPU with negligible compromise to reconstruction quality.Extensive experiments demonstrate that RegVGGT achieves state-of-the-art performance on long-horizon benchmarks across diverse FFRM prediction tasks, surpassing prior FFRM-based stream reconstruction baselines by a large margin.
CommentsECCV 2026. Corresponding author is Junjun Jiang