arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06706cs.LGcs.AI

对决世界模型:用于共模干扰项抑制的优势式动作通道

Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection

Jiazhuo Li, Yiming Fei, Zhiruo Zhou, Heikichi Hayashi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出对决世界模型,通过在潜在动态中减去动作预测的平均效应来抑制共模干扰项,无需额外机制,在多个任务中可恢复智能体自身效应,适用于各类动作条件世界模型。

中文摘要 AI 辅助

潜在世界模型通过从动作预测未来状态来进行规划,但当场景包含智能体无法控制的运动时,它们会悄然陷入动作盲:即使训练损失持续改善,不同动作的预测也变得无法区分。现有的补救措施通过重构、任务奖励或辅助目标来抑制干扰,每种措施都增加了机制或假设。我们展示了一种极简的替代方案,它借鉴了价值分解为状态基线和动作优势的对决结构:在潜在动态中,减去预测在所有动作上的平均效应,可抵消动作共有的任何部分——即干扰项所在的与动作无关的变化,从而留下一个干净、可控的通道,且无需奖励、重构或特定于干扰项的辅助损失。由于这仅为读出时的一次减法操作,因此它可原封不动地应用于任何以动作为条件的世界模型,包括冻结的预训练模型。在网格世界、具有已知因子的合成生成器、带干扰项的连续控制以及自然像素的Atari等任务中,该隔离通道在纠缠预测器失效时可恢复智能体自身的效应,且干扰项泄漏量可忽略不计;事后应用时,它能在现成模型中揭示出其原始读出缺失的动作通道,并在网格世界中转化为目标到达控制。我们证明了该抵消在有限样本下对离散和采样动作集均是精确的,并在附录中给出了其测量边界——即运动跟踪动作的干扰项,以及剩余的局限性。

英文摘要

Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quietly go action-blind: predictions for different actions become indistinguishable even as the training loss keeps improving. Existing remedies suppress this distraction with reconstruction, task reward, or auxiliary objectives, each adding machinery or assumptions. We show that a minimal alternative suffices, borrowed from the dueling decomposition of value into a state baseline and an action advantage: in latent dynamics, subtracting a prediction's mean effect over actions cancels whatever the actions share--the action-independent variation where distractors live--leaving a clean, controllable channel, with no reward, no reconstruction, and no distractor-specific auxiliary loss. Because this is only a subtraction at readout time, it applies unchanged to any action-conditioned world model, including frozen pretrained ones. Across a gridworld, synthetic generators with known factors, distracting continuous control, and natural-pixel Atari, the isolated channel recovers the agent's own effect where entangled predictors fail, with nuisance leak indistinguishable from zero; applied post hoc it surfaces an action channel in off-the-shelf models that their raw readouts miss, and it converts into goal-reaching control in the gridworld. We prove the cancellation is exact in finite samples for both discrete and sampled action sets, and we state its measured boundary--distractors whose motion tracks the action--together with the remaining limitations in the appendix.

补充信息

↑