arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OpenLongTail:长尾驾驶数据的生成式扩展

OpenLongTail: Generative Scaling of Long-Tail Driving Data

Lulin Liu, Nuo Chen, Yan Wang, Bangya Liu, Wenyan Cong, Hezhen Hu, Boris Ivanovic, Hao Wang, Ziyao Zeng, Xinyu Gong, Yang Zhou, Zixiang Xiong, Dilin Wang, Zhangyang Wang, Weisong Shi, Ruohan Zhang, Marco Pavone, Zhiwen Fan

arXiv 2607.09655首次发表:更新:

发表机构

Texas A&M University; NVIDIA; UW–Madison; UT Austin; Yale University; Adobe; Meta; University of Delaware; Stanford University(德克萨斯农工大学; 英伟达; 威斯康星大学麦迪逊分校; 德克萨斯大学奥斯汀分校; 耶鲁大学; 奥多比公司; Meta; 特拉华大学; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对长尾驾驶数据稀缺影响策略扩展的问题,提出开源生成数据引擎OpenLongTail,通过姿态外推视图合成管道及普吕克射线几何增强,合成异构数据提升闭环驾驶稳健性,验证了其多方面有效性。

AI 中文摘要

扩展稳健的驾驶策略从根本上受到策划数据集中边缘情况稀缺的限制。现实世界不断捕捉这些关键事件,但从异构源收集时,此类长尾事件仍未得到充分利用。具体而言,多样但有价值的野外长尾视频缺乏训练策略模型所需的全视图覆盖,常缺少多视图姿态或仅来自单目行车记录仪。这种模态差距阻碍了这些普遍观察结果转化为用于长尾泛化的可扩展训练数据。我们引入了OpenLongTail,一个用于在长尾事件下扩展自动驾驶策略的开源生成数据引擎。为了将异构数据源转换为对策略学习有用的视图对齐且时间连贯的多视图资产,我们开发了一个基于姿态的外推视图合成管道来生成缺失视图。我们还通过将普吕克射线几何注入可扩展生成引擎,进一步增强新生成视图的跨视图一致性和时间对齐。通过合成异构长尾数据,我们观察到在处理长尾事件时闭环驾驶稳健性有显著提高。通过测量外推视图合成和姿态指标,我们验证了OpenLongTail在视觉保真度、跨视图一致性和自我轨迹恢复方面的有效性。

英文摘要

Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.

CommentsProject page: https://openlongtail.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑