arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29772cs.RO

自我感知主动学习实现自动驾驶的持续改进

Self-Aware Active Learning Enables Continual Improvement in Autonomous Driving

  • The Hong Kong Polytechnic University(香港理工大学)
  • Adelaide University(阿德莱德大学)
  • University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

Dong Hu, Chao Huang, Carman K. M. Lee, Dimitrios Kanoulas

AI总结:

该研究提出自我感知主动学习框架SAGE,通过生成恐惧与好奇心信号调节风险,在多场景中提升自动驾驶系统鲁棒性、减少安全违规并维持任务性能,实现训练后持续改进。

AI中文摘要:

基于学习的自动驾驶(AD)系统在熟悉场景中可可靠运行,但罕见的分布偏移和长尾事件仍是突发故障的主要来源。核心局限在于,多数智能体主要从被动经验中学习,缺乏估计自身能力不足、寻求及时协助并将安全关键事件转化为针对性改进的机制。本文提出自我感知引导探索(SAGE),这是一种用于自动驾驶训练后自适应的主动学习框架。SAGE学习一个预测世界模型,生成两个在线内在信号:恐惧(fear),用于估计短程预测风险和模型不确定性;好奇心(curiosity),通过预测误差衡量新颖性。好奇心自适应校准恐惧的干预阈值,使智能体能以依赖上下文的方式调节风险。当预测恐惧超过该自适应阈值时,智能体将控制权转移给专家或 fallback 策略,并利用产生的接管轨迹进行针对性模仿学习。同时,恐惧被整合到策略优化和评估中,作为面向安全的约束,以减少自适应过程中的性能退化。我们在模拟路线转移任务、基于 Waymo 的记录驾驶场景、CARLA 遮挡危险及真实世界移动机器人导航测试中评估 SAGE。在这些设置中,SAGE 提升了新颖和安全关键场景的鲁棒性,减少了安全违规,并保持与强基线策略相当的任务性能。这些结果表明,智能体可通过估计自身能力极限、在需要时请求指导并从罕见高价值事件中选择性学习,在初始训练后实现改进。

英文摘要:

Learning-based autonomous driving (AD) systems can perform reliably in familiar conditions, yet rare distribution shifts and long-tail events remain a major source of abrupt failure. A central limitation is that most agents learn primarily from passive experience and lack mechanisms to estimate when their competence is insufficient, seek timely assistance, and convert safety-critical encounters into targeted improvement. Here we present self-aware guided exploration (SAGE), an active learning framework for post-training adaptation in AD. SAGE learns a predictive world model that generates two online intrinsic signals: fear, which estimates short-horizon predictive risk and model uncertainty, and curiosity, which measures novelty through prediction error. Curiosity adaptively calibrates the intervention threshold for fear, allowing the agent to regulate risk in a context-dependent manner. When predicted fear exceeds this adaptive threshold, the agent transfers control to an expert or fallback policy and uses the resulting takeover trajectories for focused imitation learning. In parallel, fear is integrated into policy optimization and evaluation as a safety-oriented constraint to reduce performance regressions during adaptation. We evaluate SAGE in simulated route-transfer tasks, Waymo-based logged driving scenarios, CARLA occlusion hazards, and real-world mobile robot navigation tests. Across these settings, SAGE improves robustness in novel and safety-critical scenarios, reduces safety violations, and maintains task performance comparable to strong baseline policies. These results suggest that agents can improve after initial training by estimating the limits of their competence, requesting guidance when needed, and learning selectively from rare high-value events.

↑