发表机构
University of the Chinese Academy of Sciences; International Frontier Interdisciplinary Research Institute (IFIRI), Wenzhou-Kean University; OPI Labs(中国科学院大学; 温州肯恩大学国际前沿跨学科研究院; OPI实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对潜在世界模型中预测误差低但规划不可靠的问题,提出行动-后果对齐(ACA)训练目标,通过惩罚替代行动预测误差优势,提升规划性能并减少真实目标误差,且无需额外组件或交互。
AI 中文摘要
潜在世界模型学习预测观测到的状态转移,然而仅凭低预测误差并不能保证规划的可靠性。受神经科学中自我挠痒实验的启发——该实验表明,干扰运动感觉对应关系会增加预测失配——我们考察了所学世界模型是否保持类似的行动-后果对应关系。我们的结果显示,邻近的替代行动可能获得更低的预测误差,尽管它们产生的物理结果与记录目标相距更远。当该未来状态被视为目标时,这揭示了一个具体的预测-规划失配:模型对实现目标精确度较低的行动赋予了更低的成本。为缓解这一差距,我们引入了行动-后果对齐(ACA),这是一种训练目标,通过惩罚局部搜索的替代行动相对于实际行动的预测误差优势来补充前向预测,且无需额外的模型组件或训练期间的环境交互。同样的原则也可用于指导额外的数据收集以实现自我改进。我们证明,在多种环境和评估设置下,ACA提高了规划性能并减少了真实目标误差,而ACA引导的数据收集优于随机局部采样。这些结果支持行动-后果对齐作为连接预测学习与可靠规划的实用原则。
英文摘要
Latent world models learn to predict observed transitions, yet low prediction error alone does not guarantee reliable planning. Inspired by self tickling experiments in neuroscience showing that disrupting motor sensory correspondence increases prediction mismatch, we examine whether learned world models preserve an analogous action consequence correspondence.The results show nearby alternatives can receive lower prediction errors despite producing physical outcomes farther from the recorded target. With that future treated as a goal, this reveals a concrete prediction planning mismatch: the model assigns a lower cost to an action that achieves the target less accurately. To mitigate this gap, we introduce Action Consequence Alignment (ACA), a training objective that complements forward prediction by penalizing the prediction error advantage of locally searched alternatives over factual actions without additional model components or environment interactions during training. The same principle can also guide additional data collection for self improvement. We demonstrate that across diverse environments and evaluation settings, ACA improves planning performance and reduces real goal error, while ACA guided data collection outperforms random local sampling. These results support action consequence alignment as a practical principle for bridging predictive learning and reliable planning.