arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32175cs.CVcs.AI

OneFixer:面向驾驶场景的高质量且一致的单步自回归3DGS细化方法

OneFixer: High-Quality and Consistent One-Step Autoregressive 3DGS Refinement for Driving Scenes

Boseong Jeon, Junhyeop Lee, Juhan Cha, Hayoung Kim

首次发表
浏览论文内容

中文总结 AI 辅助

OneFixer提出单步自回归视频扩散修复器,通过部署匹配的共享滚动联合学习保真度和鲁棒性,在驾驶场景中实现低延迟、高一致性的3DGS细化,显著降低碰撞率。

中文摘要 AI 辅助

自回归视频扩散是一种有前景的渲染时修复器,用于自动驾驶仿真中的3D高斯泼溅(3DGS),但部署要求低延迟下具有高视觉质量和时间一致性。这对于单步因果生成尤其困难,因为每个不完美的预测会立即成为后续帧的上下文。现有方法通过多模块的分阶段训练和滚动感知正则化来稳定滚动生成,但单步质量仍达不到部署要求。我们提出OneFixer,一种在单一任务特定适应阶段训练的单步自回归视频扩散修复器。我们的关键思想是部署匹配的共享滚动:模型自身的单步预测作为流匹配的因果上下文,使训练暴露于部署时的错误,同时同一滚动接收直接的像素空间感知监督以保留细节。由于为当前帧质量优化的预测恰好被复用为未来上下文,保真度和自回归鲁棒性被联合学习,无需双向到因果转换或师生蒸馏。OneFixer进一步利用驾驶仿真容易提供的线索,即车道几何和动态智能体状态,以提高几何保真度。在Waymo和专有驾驶场景的900帧滚动中,OneFixer在单步下实现了所有基线中最低的FVD、LPIPS和DISTS,时间一致性匹配或超过多阶段DMD流水线。在相同骨干和条件下,它匹配多阶段DMD-自强制流水线,但GPU小时数不到其一半,并持续改进超过其平台期。在闭环仿真中,OneFixer相对于原始3DGS渲染将碰撞率降低了三分之一。项目页面:此https URL

英文摘要

Autoregressive video diffusion is a promising render-time fixer for 3D Gaussian Splatting (3DGS) in autonomous-driving simulation, but deployment demands high visual quality and temporal consistency at low latency. This is especially hard for one-step causal generation, where each imperfect prediction immediately becomes context for subsequent frames. Existing approaches stabilize rollouts through staged training with multiple modules and rollout-aware regularization, yet one-step quality still falls short of what deployment requires. We introduce OneFixer, a one-step autoregressive video-diffusion fixer trained in a single task-specific adaptation stage. Our key idea is a deployment-matched shared rollout: the model's own one-step predictions serve as the causal context for flow matching, exposing training to deployment-time errors, while the same rollout receives direct pixel-space perceptual supervision to preserve fine detail. Because the predictions optimized for current-frame quality are exactly those reused as future context, fidelity and autoregressive robustness are learned jointly, without bidirectional-to-causal conversion or teacher-student distillation. OneFixer further exploits cues that driving simulation readily provides, lane geometry and dynamic-agent states, to improve geometric fidelity. On Waymo and proprietary driving scenes with 900-frame rollouts, OneFixer achieves the lowest FVD, LPIPS, and DISTS among all baselines at one step, with temporal consistency matching or exceeding multi-stage DMD pipelines. Under identical backbone and conditioning, it matches a multi-stage DMD-with-Self-Forcing pipeline in under half the GPU-hours and keeps improving beyond its plateau. In closed-loop simulation with a driving policy, OneFixer reduces the collision rate by a third relative to raw 3DGS rendering. Project page: https://onefixer-web.vercel.app/

↑