RCVLA:用于自动驾驶的4D雷达接地语义推理与轨迹仲裁
RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving
中文总结 AI 辅助
提出RCVLA框架,通过雷达-语言对齐和雷达测量细化轨迹,在OmniHD-QA上显著提升语义推理和轨迹规划性能。
中文摘要 AI 辅助
4D雷达提供了几何和运动线索,补充了视觉语义,但将其集成到视觉-语言-动作(VLA)模型中,既需要雷达-语言对齐以进行语义推理,也需要显式使用雷达测量进行轨迹细化和选择。为支持这些能力,我们构建了Cap4DR,包含86,016个雷达-图像-文本样本,用于对齐预训练;以及OmniHD-QA,包含520,161个问答对,用于场景描述、关键对象推理、占用理解和轨迹规划的指令微调。基于这些数据集,我们提出了RCVLA,一个雷达-相机VLA框架,包括雷达接地语义推理阶段(RCVLA-Sem)和轨迹仲裁阶段(RCVLA-Phys)。RCVLA-Sem在相机和雷达令牌之间执行门控双向交互,用于驾驶问答和参考轨迹生成,而辅助头提供对象和占用查询。RCVLA-Phys通过以这些查询和聚类级雷达测量为条件的截断扩散来细化参考引导的轨迹候选,然后使用雷达衍生的碰撞时间风险校准候选分数。在OmniHD-QA上,RCVLA-Sem相对于OmniDrive将CIDEr提高了9.92分,并将关键对象速度误差降低了21.9%。RCVLA-Phys相对于RCVLA-Sem将平均L2误差从0.348米降低到0.259米,平均开环碰撞率从0.576%降低到0.175%。消融研究进一步表明,语言对齐的雷达令牌改善了语义推理,而聚类级雷达测量和风险校准改善了轨迹仲裁。代码将发布。
英文摘要
4D radar provides geometric and motion cues that complement visual semantics, but integrating it into vision-language-action (VLA) models requires both radar--language alignment for semantic reasoning and explicit use of radar measurements for trajectory refinement and selection. To support these capabilities, we construct Cap4DR with 86,016 radar-image-text samples for alignment pretraining and OmniHD-QA with 520,161 question-answer pairs for instruction tuning across scene description, key-object reasoning, occupancy understanding, and trajectory planning. Building on these datasets, we propose RCVLA, a radar-camera VLA framework consisting of a radar-grounded semantic reasoning stage (RCVLA-Sem) and a trajectory arbitration stage (RCVLA-Phys). RCVLA-Sem performs gated bidirectional interaction between camera and radar tokens for driving question answering and reference trajectory generation, while auxiliary heads provide object and occupancy queries. RCVLA-Phys refines reference-guided trajectory candidates through truncated diffusion conditioned on these queries and cluster-level radar measurements, then calibrates candidate scores using radar-derived time-to-collision risk. On OmniHD-QA, RCVLA-Sem improves CIDEr by 9.92 points and reduces key-object velocity error by $21.9\%$ relative to OmniDrive. RCVLA-Phys further reduces average L2 error from $0.348$ to $0.259\,\mathrm{m}$ and average open-loop collision rate from $0.576\%$ to $0.175\%$ relative to RCVLA-Sem. Ablation studies further show that language-aligned radar tokens improve semantic reasoning, while cluster-level radar measurements and risk calibration improve trajectory arbitration. Code will be released.