发表机构
Uber AV Labs; Purdue University(优步自动驾驶实验室; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出TraVEL框架,将通用多模态嵌入模型适配到驾驶视频检索,结合轨迹监督与Group Relative Policy Optimization,在nuReasoning基准上提升了不同规模模型的运动相关检索性能。
AI 中文摘要
从大规模驾驶日志中高效检索相关片段,对数据整理、模型开发及安全分析至关重要。结构化基于规则的检索系统可明确针对驾驶事件,但通常需要专家定义的规则、辅助数据及多阶段感知流水线。多模态嵌入模型提供了更简单高效的替代方案,其用单个可搜索向量表示每个视频。然而通用模型常依赖静态场景上下文的捷径,难以区分以运动为核心的事件,如左转与右转、加速与减速。本研究探讨如何将通用多模态嵌入模型适配到驾驶视频检索任务。我们首先使用InfoNCE损失函数,在来自nuReasoning的配对片段与推理轨迹上对Qwen3-VL-Embedding进行微调。尽管该阶段大幅提升了整体检索效果,但仅靠字幕监督仍不足以实现细粒度运动理解。因此我们引入TraVEL(Trajectory-Guided Video Embedding Learning,轨迹引导视频嵌入学习),这是一种运动感知微调框架,在Group Relative Policy Optimization中以自车轨迹相似度作为奖励。轨迹仅作为特权训练监督;检索仍基于单向量视频嵌入,无需自车位姿、专家规则或辅助感知输出。我们还基于nuReasoning构建了一个驾驶视频检索基准。实验表明,TraVEL在各模型规模下均提升了以运动为核心的检索效果:相较于SFT,其在2B规模下将纵向和横向mAP分别提升9.8和4.7个点,在8B规模下对应提升7.2和1.5个点。TraVEL因此将基于物理的监督与高效的基于嵌入的检索相结合。
英文摘要
Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule-based retrieval systems can explicitly target driving events, but typically require expert-defined rules, auxiliary data, and multi-stage perception pipelines. Multimodal embedding models offer a simpler and more efficient alternative by representing each video with a single searchable vector. However, general-purpose models often rely on shortcuts from static scene context and struggle to distinguish motion-centric events, such as turning left versus right or accelerating versus decelerating. In this work, we study how to adapt a general-purpose multimodal embedding model to driving-video retrieval. We first fine-tune Qwen3-VL-Embedding on paired clips and reasoning traces from nuReasoning using an InfoNCE objective. While this stage substantially improves overall retrieval, caption supervision alone remains insufficient for fine-grained motion understanding. We therefore introduce TraVEL (Trajectory-Guided Video Embedding Learning), a motion-aware fine-tuning framework that uses ego-trajectory similarity as a reward within Group Relative Policy Optimization. Trajectories serve only as privileged training supervision; retrieval still operates on single-vector video embeddings without ego poses, expert rules, or auxiliary perception outputs. We further construct a driving-video retrieval benchmark from nuReasoning. Experiments show that TraVEL improves motion-centric retrieval across model scales: relative to SFT, it raises longitudinal and lateral mAP by 9.8 and 4.7 points at 2B, with corresponding gains of 7.2 and 1.5 points at 8B. TraVEL thus combines physically grounded supervision with efficient embedding-based search.