arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RedLight-VLA:用于驾驶策略中交通规则落地与行为强化的模型

RedLight-VLA: Models for traffic-rule grounding and behavioral emphasis in driving policies

Bala Murali Manoghar Sai Sudhakar, Sourab Bapu Sridhar, Sandipan Das, Rahul Ahuja, Meda Lazar, Ashish Garg, Pratik Likhar, Senthil Yogamani

arXiv 2608.28656首次发表:更新:

发表机构

Qualcomm Technologies, Inc.; Qualcomm Auto Ltd Sweden; Qualcomm India Private Limited; Arriver System Software S.r.l.(高通技术公司; 瑞典高通汽车有限公司; 高通印度私人有限公司; Arriver系统软件有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对行为克隆VLA驾驶策略在信号交叉口罕见规则操作上的不足,提出RedLight-VLA,通过行为重加权与并行辅助头结合,提升交通规则落地能力,在多位移指标上优于基线及单独机制。

AI 中文摘要

行为克隆的视觉-语言-动作(VLA)驾驶策略在信号交叉口的罕见规则导向操作上表现不佳。制动和起步样本对平均轨迹损失的贡献很小,而融合表征缺乏对交通信号灯和停止线状态的显式监督。我们提出RedLight-VLA,这是一种利用专家未来轨迹和自动生成的感知目标的训练目标,无需额外的人工规则标注。首先,基于轨迹的行为重加权(BR)通过旋转不变纵向动力学和尺度保持约简来强调罕见的减速和加速,当禁用该模块时,其表现恰好恢复为基线水平。其次,并行辅助(AUX)头将交通信号灯和停止线状态落地到连续的后融合规则令牌中,无需自回归语言生成或改变轨迹解码器。我们在一组精心整理的20秒序列上进行评估,预测 horizon 为5秒。受控变体共享相同的主干网络、训练数据、解码器和评估群体。与其他方面完全相同的VLA基线相比,RedLight-VLA将红灯下停止线超调从7.3%降至6.8%,将停止线速度误差降低12.7%,并将按交通信号灯切片的3秒平均位移误差(ADE)/最终位移误差(FDE)从0.274/0.964米提升至0.247/0.897米。绿灯下的误停车从3.2%升至3.9%;不过,将BR与AUX监督结合可缓解仅使用AUX时观察到的更大增幅(4.0%)。该组合模型还将非交通信号灯相关的ADE/FDE从0.268/0.956米提升至0.241/0.876米,且在所有四个切片位移指标上均优于任一单独机制。

英文摘要

Behavior-cloned Vision-Language-Action (VLA) driving policies struggle with rare rule-governed maneuvers at signalized intersections. Braking and launching examples contribute little to averaged trajectory loss, while fused representations lack explicit supervision for the governing traffic-light and stop-line state. We present RedLight-VLA, a training objective that uses expert futures and automatically generated perception targets without additional manual rule annotation. First, trajectory-derived behavioral reweighting (BR) emphasizes rare deceleration and acceleration using rotation-invariant longitudinal dynamics and a scale-preserving reduction that exactly recovers the baseline when disabled. Second, parallel auxiliary (AUX) heads ground traffic-light and stop-line state in continuous post-fusion rule tokens, without autoregressive language generation or changes to the trajectory decoder. We evaluate on a curated set of 20 s sequences with a 5 s prediction horizon. Controlled variants share the same backbone, training data, decoder, and evaluation population. Against an otherwise identical VLA baseline, RedLight-VLA reduces red-light stop-line overshoot from 7.3% to6.8%, reduces stop-line velocity error by 12.7%, and improves 3 s trafficlight-sliced ADE/FDE from 0.274/0.964 m to 0.247/0.897 m. Green-light false stops increase from 3.2% to 3.9%; however, combining BR with AUX supervision mitigates the larger increase observed for AUX alone (4.0%). The combined model also improves non-traffic-light ADE/FDE from 0.268/0.956 m to 0.241/0.876 m and outperforms either mechanism alone on all four sliced displacement measures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑