arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越文本思维链:自动驾驶中基于动作的推理综述

Beyond Textual Chain-of-Thought: A Survey on Action-Grounded Reasoning in Autonomous Driving

Zhengxu Tang, Xiaozhou Zhang, Guofeng Cui, Ziyu Gong, Zi Wang, Yunfei Shi, Ruifeng Deng, Chengzhi Qi, Ke Chen, Sachin Patil, Tianjun Xiao, Langechuan Liu, Pichao Wang

arXiv 2609.01659首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述调研171篇文献,提出以表示为中心的分类法,将自动驾驶基于动作的推理方法分为四类13个子类型,指出其前沿为可关联现实世界、耦合实时动作且可安全验证的中间表示。

AI 中文摘要

思维链(CoT)推理通过在生成答案前引出中间步骤来驱动生成式模型。在自动驾驶中,答案是连续动作,因此其推理必须与物理世界共享相同的时空结构。本综述研究了从文本CoT到基于动作的推理的转变,共调研171篇论文,其中包括130篇方法论文和41篇基准、数据集、综述及分析论文。我们提出了以表示为中心的分类法,将中间状态的形式作为组织轴,将130种方法系统化为四类:基于语言的、视觉空间的、潜在动态的和外化的推理,进一步细分为与不同感兴趣区域相关的13个子类型。我们的综合分析表明,自动驾驶智能体推理的前沿在于能与现实世界建立关联、与实时动作耦合并能在安全关键系统下验证的中间表示。项目页面:this https URL。

英文摘要

Chain-of-thought (CoT) reasoning powers generative models by eliciting intermediate steps before producing an answer. In autonomous driving, the answer is a continuous action. Thus its reasoning must share the same spatiotemporal structure as the physical world. This survey studies the resulting shift from textual CoT to action-grounded reasoning. Surveying 171 papers, including 130 method papers and 41 benchmarks, datasets, surveys, and analysis papers, we propose a representation-centered taxonomy that treats the form of the intermediate state as the organizing axis. We systematize the 130 methods into four categories: language-based, visual-spatial, latent-dynamic, and externalized reasoning, further divided into 13 subtypes tied to distinct regions of interests. Our synthesis shows that the open frontier of reasoning in driving agents lies in intermediate representations that can be grounded in the real world, coupled to real-time action, and verified under safety-critical systems. Project page: https://github.com/tangzhengxu/awesome-av-cot.

CommentsAccepted by EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑