arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过后见之明与预见之明驾驶:面向自动驾驶的层次化记忆上的工具接地协同推理

Drive by Hindsight and Foresight: Tool-Grounded Synergistic Reasoning over Hierarchical Memory for Autonomous Driving

Baojie Chen, Zijun Jia, Jing Zhong

arXiv 2609.08217首次发表:更新:

发表机构

Beihang University; Tsinghua University(北京航空航天大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出首个将层次化记忆与主动工具调用结合的闭环推理框架,通过短期与长期记忆协同,提升自动驾驶推理和决策,7B模型在DriveLMM-o1上显著超越基线。

AI 中文摘要

视觉语言模型(VLMs)在自动驾驶领域展现出潜力,但仍存在幻觉、时空感知能力弱和泛化能力有限等问题。近期方法通过思维链(CoT)解释、检索增强生成或静态注入工具输出来改进推理和决策。尽管这些机制丰富了上下文,但模型既不能主动感知场景信息,也不能在回答后积累经验。为克服这些局限,我们提出了据我们所知首个将层次化记忆与主动工具调用紧密耦合于闭环推理循环中的协同框架。我们的贡献有三方面:(i)层次化驾驶记忆:场景级短期记忆维持动态场景状态,不断演化的长期记忆检索可复用经验和工具策略。(ii)记忆-工具协同推理框架:在场景状态和检索经验的引导下,模型在推理时自适应调用工具以优化推理,并在离线时将可复用经验整合到长期记忆池中。(iii)数据生成与两阶段训练流程:利用多步教师展开构建的验证过的记忆-工具轨迹,通过监督微调(SFT)和组相对策略优化(GRPO)进行训练。我们的7B模型在DriveLMM-o1上达到80.03的总体推理分数和79.09%的多选题(MCQ)准确率,比最强基线高出7.74个MCQ点,并在多个基准上表现出强大的泛化能力。值得注意的是,消融和分析研究验证了每个组件的有效性,并进一步揭示了层次化记忆的互补作用。短期记忆增强了时空理解,使STSBench准确率提高24.2个百分点,而离线长期记忆整合在全部参数冻结的情况下额外带来3.57个MCQ点的提升,展示了通过积累驾驶经验实现持续自我进化的能力。

英文摘要

VLMs have shown promise for autonomous driving, yet still suffer from hallucination, weak spatio-temporal perception, and limited generalization. Recent methods improve reasoning and decision-making through CoT explanations, retrieval-augmented generation or the static injection of tool outputs. Although these mechanisms enrich the context, the model neither proactively perceives scene information nor accumulates experience after answering. To overcome these limitations, we present, to our knowledge, the first synergistic framework that tightly couples hierarchical memory with proactive tool invocation in a closed reasoning loop. Our contributions are threefold. (i) Hierarchical Driving Memory: a scene-level short-term memory maintains the dynamic scene state, and an evolving long-term memory retrieves reusable experience and tool strategies. (ii) Memory-Tool Synergistic Reasoning Framework: guided by the scene state and retrieved experience, the model adaptively invokes tools to refine its reasoning at inference time and consolidates reusable experience into a long-term memory pool offline. (iii) Data Generation and Two-stage Training Pipeline: verified memory-tool trajectories built by multi-step teacher rollout are used to train with SFT and GRPO. Our 7B model reaches an overall reasoning score of 80.03 and MCQ accuracy of 79.09% on DriveLMM-o1, surpassing the strongest baseline by 7.74 MCQ points and generalizes strongly across benchmarks. Notably, ablation and analysis studies validate the effectiveness of each component and further reveal the complementary roles of hierarchical memory. Short-term memory strengthens spatio-temporal understanding, improving STSBench accuracy by 24.2 points, while offline long-term memory consolidation yields an additional 3.57-point MCQ gain with all parameters frozen, demonstrating continual self-evolution through accumulated driving experience.

Comments19 pages, 7 figures, 5 tables. Includes appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑