arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Dreamer-SAC:用于样本高效自动驾驶的潜世界模型离线策略学习

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

Jiazhuo Li, Linjiang Cao, Qi Liu, Xi Xiong

arXiv 2608.10386首次发表:更新:

发表机构

Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Dreamer-SAC框架,结合循环状态空间世界模型与离线策略SAC算法,在自动驾驶场景中优于DreamerV3、SAC等基线,且所需真实环境交互更少。

AI 中文摘要

样本高效的自动驾驶强化学习常受限于数据效率与模型偏差之间的权衡。尽管世界模型可降低对高成本环境交互的依赖,但在学习到的动力学上进行策略优化仍对预测误差敏感。本文提出Dreamer-SAC框架,它将循环状态空间世界模型与直接在潜空间中训练的离线策略软 Actor-Critic(SAC)算法相结合。该框架同时使用真实交互和短视界生成轨迹,并结合n步目标估计与多目标监督进行训练。在包含驾驶效率与安全目标的自动驾驶场景中评估时,所提框架始终优于代表性强化学习基线,包括DreamerV3、SAC和PPO,且在大幅减少真实环境交互的同时实现了性能提升。实验揭示了回退视界与策略性能之间的倒U型关系,其中短视界潜回退在额外训练信号与累积模型偏差之间实现了最佳权衡。此外,n步目标估计在利用预测经验进行价值学习方面比单步时间差分目标更有效。

英文摘要

Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world models reduce the reliance on costly environment interactions, policy optimization over learned dynamics remains sensitive to prediction errors. This paper proposes the Dreamer-SAC framework, which integrates a recurrent state-space world model with an off-policy soft actor-critic algorithm trained directly in latent space. The framework uses a combination of real interactions and short-horizon generated trajectories with n-step target estimation and multi-objective supervision. Evaluated in autonomous driving scenarios with objectives encompassing driving efficiency and safety, the proposed framework consistently outperforms representative reinforcement learning baselines, including DreamerV3, SAC, and PPO, while achieving improved performance with substantially fewer real environment interactions. Experiments reveal an inverted-U relationship between rollout horizon and policy performance, where short-horizon latent rollouts achieve the best trade-off between additional training signals and accumulated model bias. Furthermore, n-step target estimation demonstrates more effectiveness over one-step temporal-difference targets in exploiting predicted experience for value learning.

Comments13 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑