arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Adaptive-WAM:基于质量引导的中间视频扩散特征早期退出规划

Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features

Sining Ang, Yuguang Yang, Yan Wang

arXiv 2608.06008首次发表:更新:

AI 中文总结

Adaptive-WAM是基于Wan2.2-5B主干的质量感知多出口规划器,通过动态分配主干深度减少计算量,在NAVSIM、nuScenes等数据集上取得优异规划性能,延迟显著降低。

AI 中文摘要

大型视频扩散模型为自动驾驶提供了丰富的时空先验,但现有的世界-动作模型通常继承了迭代未来视频生成的成本,而部署仅需要自车轨迹。我们提出一个更基础的问题:为做出可靠的驾驶决策,必须执行视频扩散模型的多少部分?通过对视频去噪时间步和扩散Transformer(DiT)深度的控制研究,我们发现规划性能对所测试的视频噪声水平基本不敏感,而强大的轨迹已可从中间层解码。基于此观察,我们引入Adaptive-WAM,一种基于Wan2.2-5B主干构建的质量感知多出口规划器。轨迹扩散头被附加到选定的DiT块,轻量级轨迹质量评分器会在目前解码的最佳轨迹满足质量阈值时终止推理;否则,计算从缓存的隐藏状态继续到更深的出口。因此,部署的规划器避免了未来视频合成所需的迭代无分类器去噪循环和VAE解码,同时根据轨迹质量动态分配主干深度。在NAVSIM上,自适应单轨迹规划器达到90.8 PDMS;单独的固定出口变体使用64个提议达到92.6 PDMS。其在NAVSIM v2上进一步获得89.9 EPDMS,在对比的前视图视频世界模型规划器中取得报告的最佳结果。无目标域微调的情况下,Adaptive-WAM以0.88米的平均L2误差和0.08%的碰撞率迁移到nuScenes。在A100上,自适应路由将PDMS从90.62提升至90.79,同时平均端到端规划延迟为170毫秒,比190毫秒的固定第15块规划器低约10%,比320毫秒的固定全深度规划器低47%。代码将发布。

英文摘要

Large video diffusion models provide rich spatiotemporal priors for autonomous driving, but existing world-action models often inherit the cost of iterative future-video generation even though deployment only requires an ego trajectory. We ask a more basic question: how much of a video diffusion model must be executed to make a reliable driving decision? Through a controlled study of video denoising timesteps and Diffusion Transformer (DiT) depth, we find that planning performance is largely insensitive to the tested video-noise levels, whereas strong trajectories can already be decoded from intermediate layers. Based on this observation, we introduce Adaptive-WAM, a quality-aware multi-exit planner built on a Wan2.2-5B backbone. Trajectory diffusion heads are attached to selected DiT blocks, and a lightweight trajectory-quality scorer terminates inference once the best trajectory decoded so far satisfies a quality threshold; otherwise, computation continues from the cached hidden state to a deeper exit. The deployed planner therefore avoids the iterative classifier-free denoising loop and VAE decoding required for future-video synthesis, while dynamically allocating backbone depth according to trajectory quality. On NAVSIM, the adaptive single-trajectory planner achieves 90.8 PDMS; a separate fixed-exit variant reaches 92.6 PDMS with 64 proposals. It further obtains 89.9 EPDMS on NAVSIM v2, yielding the best reported results among the compared front-view video world-model planners. Without target-domain fine-tuning, Adaptive-WAM transfers to nuScenes with 0.88 m average L2 error and a 0.08\% collision rate. On an A100, adaptive routing improves PDMS from 90.62 to 90.79 while averaging 170 ms end-to-end planning latency, approximately 10\% below the 190 ms fixed block-15 planner and 47\% below the 320 ms fixed full-depth planner. Code will be released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑