发表机构
Peking University; Xiaomi EV(北京大学; 小米电动汽车)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自动驾驶视觉语言模型在道路施工区域的难题,提出WorkDrive框架,通过自动化多任务感知管道提取场景事实,经监督微调与强化学习,在ROADWork数据集上降低轨迹平均位移误差,实现渐进改进。
AI 中文摘要
自动驾驶视觉语言模型(VLMs)在道路施工区域面临困难,熟悉的视觉线索改变或缺失,临时设备重新定义可行驶通道。VLMs能检测物体,但缺乏明确指导。我们提出WorkDrive框架,构建基于感知的施工区域因果推理并与轨迹预测对齐。自动化多任务感知管道提取结构化场景事实注入因果链注释管道,推理标签用于监督微调及强化学习。在最大公共施工区域数据集ROADWork上,所提道路施工因果链使轨迹平均位移误差(ADE)降低9.0%,基于一致性的GRPO再降低3.0%。代码和数据将公开发布。
英文摘要
Autonomous driving vision-language models (VLMs) struggle in roadwork zones, where familiar visual cues such as lane markings and permanent signs are altered or absent, and temporary devices such as cones and barriers redefine the drivable corridor. VLMs can detect these objects, but without explicit guidance they anchor their reasoning on familiar elements from pre-training and fail to connect work-zone observations to correct planning decisions. We propose WorkDrive, a framework that constructs perception-grounded causal reasoning for work zones and aligns it with trajectory prediction. An automated multitask perception pipeline extracts structured scene facts and injects them into a Chain-of-Causation (CoC) annotation pipeline, redirecting the annotator's attention to domain-specific elements. The resulting reasoning labels are used for supervised fine-tuning, followed by reinforcement learning with a single reward: consistency between lateral meta-actions and the predicted trajectory. On ROADWork, the largest public work-zone dataset, the proposed roadwork CoC reduces trajectory average displacement error (ADE) by 9.0\%, and consistency-based GRPO yields a further 3.0\%, achieving progressive improvement over the trajectory-only baseline. Code and data will be publicly released.