arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15142cs.AIcs.LG

《雅达利乒乓球游戏中用于世界模型的概念引导空间正则化》

Improving Weak World Models Behind Strong Agents in Atari Pong

Yukuan Lu, Zaishuo Xia, Weyl Lu, Yubei Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究雅达利乒乓球游戏中五个视觉世界模型智能体,发现其存在诸多问题。提出概念引导空间正则化(CGSReg),实验表明该方法在部分模型中改善了闭环展开和像素空间零样本MBRL,但不能解决所有世界模型瓶颈。

中文摘要 AI 辅助

世界模型通常作为基于模型的强化学习(MBRL)系统的组件进行评估,而世界模型本身很少被单独研究。我们研究了雅达利乒乓球游戏中的五个代表性视觉世界模型智能体:DreamerV3、DIAMOND、TWISTER、Simulus和STORM。在重现它们的训练流程并匹配报告的智能体性能后,冻结学习到的世界模型,用闭环展开诊断进行评估:一个与相应MBRL智能体分开训练的策略与每个冻结模型交互,并检查生成的视频轨迹中的视觉和动态错误。所有五个模型的展开都包含明显失败,包括球消失、球运动不正确和无效的球与球拍交互。除视觉轨迹外,还用像素空间零样本MBRL进一步评估,所有五个模型产生的策略都明显逊于相应原始MBRL训练流程产生的策略。我们假设对任务关键概念(如乒乓球游戏中的球)建模不足可能导致这些失败。因此提出概念引导空间正则化(CGSReg),一种应用于分割概念区域的辅助像素重建损失。实验表明,CGSReg在DreamerV3、DIAMOND和TWISTER中改善了闭环展开和像素空间零样本MBRL。其效果因其余模型和评估指标而异,表明仅CGSReg不能解决所有世界模型瓶颈。

英文摘要

Strong world-model agents frequently contain weak world models. We study this agent-world-model gap by reproducing five visual world-model agents in Atari Pong: DreamerV3, DIAMOND, TWISTER, Simulus, and STORM, with performance comparable to the reported results, and independently evaluating their frozen world models. First, closed-loop rollout diagnosis qualitatively inspects visual trajectories generated by each frozen model under an independently trained policy. All five models exhibit clear visual or dynamical failures, including ball disappearance, incorrect motion, and invalid ball-paddle interactions. Second, under native zero-shot model-based reinforcement learning (MBRL), a new policy is trained entirely within the frozen model from scratch using the agent's native RL procedure, without real-environment training. When evaluated in the real environment, these policies substantially underperform the reproduced agents: DreamerV3 (-5.5 to -20.9), DIAMOND (19.7 to -9.6), TWISTER (17.7 to -13.3), Simulus (20.8 to -11.6), and STORM (18.7 to -21.0), where -21 is the minimum Pong return. This gap also extends broadly across Atari100K. Motivated by the ball-related rollout failures in Pong, we propose Concept-Guided Spatial Regularization (CGSReg), an auxiliary reconstruction loss on task-critical concept regions. We evaluate it under a more challenging pixel-space zero-shot MBRL setting, where policies learn directly from images generated by the frozen world model. Ball-region CGSReg improves pixel-space zero-shot MBRL in DreamerV3, DIAMOND, TWISTER, and Simulus, and also improves closed-loop rollouts in the first three; STORM shows no clear improvement.

发表机构

  • UC Davis(加州大学戴维斯分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑