作为对手的世界模型:用于鲁棒运动规划的多智能体自我博弈微调
World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning
浏览论文内容
中文总结 AI 辅助
研究在密集交通中自动驾驶车辆鲁棒运动规划问题,提出对抗世界建模(AWM)框架,通过多智能体自我博弈微调,引入解耦求解器,经实验验证该方法能生成可转移对抗交互,使规划器在多种场景有竞争力。
中文摘要 AI 辅助
在密集交通中进行鲁棒运动规划,要求自动驾驶车辆在自然驾驶数据中未充分体现的罕见且安全关键场景中进行交互。尽管对抗训练提供了可行方案,但现有方法常依赖外部场景生成器、启发式扰动或大量模拟器展开,难以与现代自回归规划器集成。本文将对抗鲁棒规划器学习视为约束极小极大博弈,提出对抗世界建模(AWM),这是一个有理论基础的多智能体自我博弈微调框架。因求解精确博弈难处理,AWM引入原则性解耦求解器。在内层最小化中,规划器的预测世界模型转化为角色条件对手,通过反事实信用分配学习稀疏、场景自适应攻击联盟。在外层最大化中,自我规划器针对冻结的AWM优化遗憾感知鲁棒最佳响应,利用尾部风险加权和参考锚定信任区域改善困难情况恢复,同时保持标称驾驶行为。在nuPlan和InterPlan基准测试上的实验表明,该方法生成可转移的对抗交互,并产生一个鲁棒规划器,在标称和高度交互的长尾场景中均实现有竞争力的闭环性能。理论分析证明了解耦求解器和主要优化组件的合理性。
英文摘要
Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepresented in naturalistic driving data. Although adversarial training offers a feasible solution, existing methods often rely on external scenario generators, heuristic perturbations, or simulator-heavy rollouts, which makes them difficult to integrate with modern autoregressive planners. Here, we cast adversarially robust planner learning as a constrained min-max game and propose Adversarial World Modeling (AWM), a theoretically grounded multi-agent self-play fine-tuning framework. Since solving the exact game is intractable, AWM introduces a principled decoupled solver. In the inner minimization, the planner's predictive world model is converted into a role-conditioned adversary that learns sparse, scene-adaptive attack coalitions via counterfactual credit assignment. In the outer maximization, the ego planner optimizes a regret-aware robust best response against the frozen AWM, utilizing tail-risk weighting and reference-anchored trust regions to improve hard-case recovery while preserving nominal driving behavior. Experiments on the nuPlan and InterPlan benchmarks demonstrate that our method generates transferable adversarial interactions and yields a robust planner that achieves competitive closed-loop performance in both nominal and highly interactive long-tail scenarios. Theoretical analysis justifies the decoupled solver and the main optimization components.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
- Tongji University(同济大学)
- McGill University(麦吉尔大学)
- Mila-Quebec AI Institute(米拉-魁北克人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。