arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17739cs.MA

面向协同混合交通控制的、融入物理信息世界模型的离线多智能体强化学习

Offline Multi-Agent Reinforcement Learning with a Physics-Informed World Model for Cooperative Mixed Traffic Control

Lu Liu, Chi Xie, Xi Xiong

中文总结 AI 辅助

本研究针对混合交通中部分可观测高速瓶颈的网联自动驾驶车辆协同控制问题,提出融入物理信息世界模型的离线多智能体强化学习框架,经SUMO实验验证可提升状态重构与预测准确性。

中文摘要 AI 辅助

本研究针对混合交通中部分可观测高速公路瓶颈处的网联自动驾驶车辆(CAV)开展协同控制研究,旨在无需依赖完整全局交通状态或在线试错的情况下缓解拥堵。我们提出一种基于物理信息世界模型的离线多智能体强化学习框架,该框架可从CAV的局部观测-动作历史中重构出具有物理可解释性的全局交通状态,耦合的宏观-微观交通动力学提供基于物理的监督。概率集成世界模型学习交通状态转移与系统奖励,模型分歧量化认知不确定性。随后采用带悲观奖励与不确定性驱动截断的多步想象回滚进行离线策略学习。在基于SUMO的匝道瓶颈处开展的实验使用了约1×10^6条离线转移数据,结果显示物理监督提升了状态重构与世界模型预测的准确性。

英文摘要

This study investigates cooperative control of connected and automated vehicles (CAVs) at partially observable highway bottlenecks in mixed traffic, aiming to mitigate congestion without relying on complete global traffic states or online trial-and-error. We propose a physics-informed world model-based offline multi-agent reinforcement learning framework that reconstructs a physically interpretable global traffic state from local CAV observation-action histories, with coupled macroscopic-microscopic traffic dynamics providing physics-based supervision. A probabilistic ensemble world model learns traffic-state transitions and system rewards, while model disagreement quantifies epistemic uncertainty. Multi-step imagined rollouts with pessimistic rewards and uncertainty-driven truncation are then used for offline policy learning. Experiments in a SUMO-based on-ramp bottleneck using approximately $1\times10^6$ offline transitions show that physics supervision improves state reconstruction and world-model prediction accuracy.

↑