arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

5500小时驾驶数据能带来多大提升?视频扩散模型的缩放定律分析

How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models

Victor Besnier, Anh-Quan Cao, Elias Ramzi, Spyros Gidaris, Tuan-Hung Vu, Andrei Bursuc, Eloi Zablocki, Matthieu Cord

arXiv 2608.28404首次发表:更新:

发表机构

Sorbonne Université(索邦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对自动驾驶视频生成,通过对100万至90亿参数的视频扩散模型开展缩放定律分析,明确训练时长、模型规模、数据量的影响,训练出的90亿参数模型在nuScenes上达到开源最优。

AI 中文摘要

自动驾驶的视频生成无法遵循网络规模的路线:驾驶数据采集成本高昂,受隐私要求限制,无法随意抓取,因此模型必须充分利用固定的数据集。我们对从零开始在驾驶数据上训练的视频扩散模型开展了系统的缩放定律研究:构建了参数规模从100万到90亿的一系列模型,在最多5500小时的驾驶数据上以不同的曝光度进行训练。验证损失在模型规模和训练曝光度两方面均遵循一致的幂律,回答了影响训练预算的关键问题:计算资源应更用于延长训练时间还是扩大模型规模,以及是否需要更多数据。损失随训练曝光度的提升速度远快于随模型规模的提升速度,这使得在计算资源有限的情况下,延长训练时间是改进固定模型的最有效方式。然而,更大的模型仍能达到更低的渐近损失,因此当有足够的计算资源和数据时,计算最优缩放仍倾向于扩大模型规模。基于这些定律,我们训练了一个90亿参数的模型,据我们所知,这是首个从零开始在驾驶数据上训练的最大视频扩散模型:在nuScenes基准上,它创下了驾驶视频生成的新开源最优水平。我们的代码和预训练模型可在该https网址获取。NATIX正在分阶段发布底层驾驶数据。

英文摘要

Video generation for autonomous driving cannot follow the web-scale route: driving data is expensive to collect, bound by privacy requirements, and cannot be scraped at will, so models must make the most of a fixed corpus. We present a systematic scaling-law study of video diffusion models trained from scratch on driving data: a family of models from 1M to 9B parameters, trained at different exposures on up to 5,500 hours of driving. Validation loss follows consistent power laws in both model size and training exposure, answering the questions that shape a training budget: whether compute is better spent on longer training or on a larger model, and whether more data is needed. Loss improves much faster with training exposure than with model size, making longer training the most effective way to improve a fixed model under limited compute. However, larger models continue to achieve lower asymptotic loss, so compute-optimal scaling still favors increasing model size when sufficient compute and data are available. Guided by these laws, we train a 9B-parameter model, to our knowledge the largest video diffusion model trained from scratch on driving data: it sets a new open-source state of the art for driving video generation, as measured on nuScenes. Our code and pretrained models are available at https://github.com/valeoai/VATIX. NATIX is separately releasing the underlying driving data in stages.

Journal ref[Archival Track] ECCV 2026 DriveX - 6th Workshop on Foundation Models for Autonomous Driving

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑