发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对离线多智能体强化学习中教师模型生成跨模式样本并传播误差的问题,提出MoSDOT方法,通过半离散最优传输将噪声分配到单一模式,提升端点质量与路由一致性。
AI 中文摘要
离线多智能体强化学习(MARL)日益依赖生成式策略来建模多模态的联合行为,通常是在集中训练与分散执行(CTDE)框架下,将集中式教师模型蒸馏为分散的单步执行智能体。我们识别出教师训练阶段的一个失败模式:基于标准流的教师模型将噪声与回放目标独立配对,因此相近的噪声样本可能被导向相互冲突的协调模式。随后,教师模型会在有效模式之间产生样本,并且由于蒸馏损失将每个局部智能体回归到教师输出在给定局部输入下的条件均值,该误差不会被吸收,而是被传播给学生模型。为消除这一教师侧伪影,我们提出模式支持半离散最优传输(MoSDOT),该方法将多模态回放概括为具有规定容量的有限模式支持,并在教师训练之前使用条件半离散最优传输将每个噪声样本分配到单一模式。我们还研究了一种共享随机性变体,该变体在执行时使用共享噪声分量,以暴露严格乘积执行所固有的残余差距。在受控诊断和离线MARL基准测试上,MoSDOT提升了端点质量和路由一致性,尤其是在展现多模态联合行为的数据集上。
英文摘要
Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher training stage: standard flow-based teachers pair noise with replay targets independently, so nearby noise samples can be routed toward conflicting coordination modes. The teacher then produces samples between valid modes, and because the distillation loss regresses each local actor onto the conditional mean of the teacher's output given local input, this error is not absorbed but propagated to the student. To remove this teacher-side artifact, we propose Mode-Support Semi-Discrete Optimal Transport (MoSDOT), which summarizes multimodal replay into a finite mode support with prescribed capacities and uses conditional semi-discrete optimal transport to assign each noise sample to a single mode before teacher training. We additionally study a shared-randomness variant that uses a shared noise component at execution to expose the residual gap intrinsic to strict-product execution. On controlled diagnostics and offline MARL benchmarks, MoSDOT improves endpoint quality and routing consistency, particularly on datasets exhibiting multimodal joint behavior.
CommentsNeurIPS 2026, Project page: https://alex6095.github.io/mosdot/