FocusDrive:基于视觉焦点的自动驾驶推理
FocusDrive: Reasoning with Visual Focus for Autonomous Driving
浏览论文内容
中文总结 AI 辅助
FocusDrive提出结构化多模态推理框架,通过显式视觉焦点连接场景理解与驾驶行动,在W3DA和NAVSIM上实现优于文本思维链的注视预测和端到端规划性能。
中文摘要 AI 辅助
驾驶决策既取决于关注何处,也取决于如何对所见的景象采取行动。有效的驾驶推理必须确定哪些物体重要、它们位于何处,以及它们如何为预期的行动提供信息。基于文本的推理可以描述驾驶响应,但将其与特定视觉证据的对应关系隐含化。视觉焦点通过识别场景中重要的内容及其位置,为这种联系提供了具体的起点。我们提出了FocusDrive,一个结构化的多模态推理框架,围绕显式视觉焦点组织端到端规划。它将决策相关对象的描述与图像补丁引用配对,将显式视觉焦点引入生成驾驶计划和轨迹的推理中。我们首先通过驾驶员注视预测评估这种焦点表示,然后利用现有NAVSIM训练场景中的驾驶焦点注释研究其在规划推理中的作用。在W3DA和NAVSIM上的实验展示了具有竞争力的注视预测和端到端规划性能,FocusDrive优于基于文本的思维链。这些结果支持视觉焦点作为场景理解与驾驶行动之间的有效桥梁。
英文摘要
Driving decisions depend on both where to focus and how to act on what is seen. Effective driving reasoning must establish which objects matter, where they are, and how they inform the intended action. Text-based rationales can describe a driving response while leaving its correspondence to specific visual evidence implicit. Visual focus provides a concrete starting point for this connection by identifying what matters in the scene and where it is. We propose FocusDrive, a structured multimodal reasoning framework that organizes end-to-end planning around explicit visual focus. It pairs descriptions of decision-relevant objects with image-patch references, bringing explicit visual focus into the reasoning that generates driving plans and trajectories. We first assess this focus representation through driver gaze prediction, then investigate its role in planning reasoning using driving-focus annotations within existing NAVSIM training scenes. Experiments on W3DA and NAVSIM demonstrate competitive gaze prediction and end-to-end planning performance, with FocusDrive improving over text-based chain-of-thought. These results support visual focus as an effective link between scene understanding and driving action.