arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理以规范:用于交通规则理解的思维链

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

Yueru Luo, Xu Yan, Changqing Zhou, Yiming Yang, Chao Zhan, Shuqi Mei, Chao Zheng, Zhen Li

arXiv 2607.24199首次发表:更新:

发表机构

CUHK-SZ, Shenzhen-FNii; HKUST(GZ); Tencent T Lab(香港中文大学(深圳),深圳高等金融研究院; 香港科技大学(广州); 腾讯T实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对交通规则理解这一自动驾驶关键挑战,提出为视觉语言模型配备思维链能力的框架,设计整理管道,采用两阶段训练方案,显著提升了可解释性和准确性,建立首个基于推理的自动驾驶框架。

AI 中文摘要

理解和遵守交通规则是自动驾驶安全的关键要求,但由于交通标志的多样性和上下文依赖性,这仍然具有挑战性。规则理解不是简单的识别任务,而是推理问题。MapDR提供了细粒度注释。现有方法多将其视为直接序列预测,忽略了连接标志语义和地图结构的潜在推理。为此,我们明确将推理纳入任务,提出为视觉语言模型配备思维链能力的框架。设计了可扩展的思维链整理管道,采用两阶段训练方案。在MapDR上的大量实验表明,我们的方法显著提高了可解释性和准确性,建立了首个基于推理的自动驾驶框架。

英文摘要

Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversity and context dependence of traffic signage. Importantly, regulation understanding is not a simple recognition task, but a reasoning problem: whether a rule applies depends on interpreting the sign in relation to the spatial layout of lanes and scene context. To support such reasoning, MapDR provide fine-grained annotations that link each traffic sign's regulatory rules to the specific lanes they govern. Existing methods, however, largely treat this as direct sequence prediction, ignoring the underlying reasoning that connects sign semantics and map structure. To address this limitation, we explicitly incorporate reasoning into this task and propose a framework that equips vision-language models (VLMs) with chain-of-thought (CoT) capabilities. We first design a scalable CoT curation pipeline that bootstraps rationales from a strong LLM through a two-round strategy and employs a VLM-based verifier to filter out incorrect cases, yielding a high-quality set of (CoT, answer) pairs. Building on this foundation, we adopt a two-stage training scheme: supervised fine-tuning (SFT) to teach rationale-to-answer generation, followed by GRPO reinforcement learning with answer-grounded, fine-grained rewards to further improve final answer accuracy. Extensive experiments on MapDR show that our approach significantly improves both interpretability and accuracy, establishing the first reasoning-based framework for regulation-aware autonomous driving.

CommentsTechnical Report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑