arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越策略支持:交互约束的离线强化学习用于自动驾驶

Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving

Mahmoud Selim, Cristina Cipriani, Karl Henrik Johansson

arXiv 2610.09763首次发表:更新:

发表机构

KTH Royal Institute of Technology(瑞典皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自动驾驶离线强化学习中的交互分布偏移问题,提出交互约束驾驶策略(ICDP),通过对比密度比估计分解联合支持退化,无需联合密度建模或模拟器 rollout,在 nuPlan、Interplan 和真实卡车实验中提升交互关键场景性能。

AI 中文摘要

离线强化学习能够从固定数据集中实现奖励驱动的策略改进,而无需在线探索,这使得它在安全关键领域特别具有吸引力。然而,一个核心挑战是分布偏移:策略优化可能倾向于选择那些在离线数据中支持较弱的动作,导致价值估计不可靠。现有方法主要在其自身的动作空间中控制这种偏移。在诸如自动驾驶之类的交互环境中,这可能是不够的:候选的自车轨迹在边际行为分布下可能得到良好支持,但在与日志交互中观察到的周围智能体行为联合起来时,支持可能较差。我们将这种交互支持的退化称为交互分布偏移(IDS),并引入交互约束驾驶策略(ICDP),这是一个显式控制交互级分布偏移的离线强化学习框架。从自车和周围智能体未来的联合数据分布出发,我们展示了联合支持的退化可以精确分解为自车支持分量和残差交互支持分量。我们通过对比密度比估计来恢复后者,从而在策略优化过程中,无需显式联合密度建模、周围智能体预测,也无需在反应式模拟器或学习的世界模型中进行 rollout,即可隔离交互兼容性。在 nuPlan、Interplan 和真实世界卡车实验上的闭环评估表明,ICDP 抑制了高价值但交互不支持的轨迹选择,并在交互关键的驾驶场景中提升了性能。项目网页:此 https URL

英文摘要

Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offline data, rendering value estimates unreliable. Existing approaches primarily control this shift in the policy's own action space. In interactive environments such as autonomous driving, this can be insufficient: a candidate ego trajectory may remain well supported under the marginal behavior distribution while being poorly supported jointly with the surrounding-agent behavior observed in the logged interaction. We refer to this degradation in interaction support as \emph{interaction distribution shift} (IDS), and introduce \emph{Interaction-Constrained Drive Policy} (ICDP), an offline reinforcement learning framework that explicitly controls interaction-level distribution shift. Starting from the joint data distribution over ego and surrounding-agent futures, we show that joint-support degradation decomposes exactly into an ego-support component and a residual interaction-support component. We recover the latter through contrastive density-ratio estimation, isolating interaction compatibility without explicit joint-density modeling, surrounding-agent prediction, or rollouts in reactive simulators or learned world models during policy optimization. Closed-loop evaluations on nuPlan, Interplan and real-world truck experiments show that ICDP suppresses high-value yet interaction-unsupported trajectory selections and improves performance in interaction-critical driving scenarios. Project webpage: https://mahmoud-selim.github.io/ICDP/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑