arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

世界模型策略学习与模仿世界-动作模型之间的能力分离

On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models

Yang Yu

arXiv 2608.22197首次发表:更新:

发表机构

Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比了直接行为克隆策略等三类策略,明确世界-动作模型学习与直接行为克隆的能力差异,指出观测演示无法识别动作效果,干预可实现零遗憾值。

AI 中文摘要

世界-动作模型会预测未来结果,进而推断出对应的动作。尽管这种因子分解能够提升表示学习效果与数据效率,但当二者均基于相同的观测演示进行训练时,其是否能提供比直接行为克隆更强的控制能力尚不明确。我们对比了直接行为克隆策略、经模仿训练的世界-动作策略,以及通过动作条件世界模型优化得到的策略。在控制器类别层面,所有世界-动作策略均可被扁平化转化为具有相同闭环轨迹分布的直接随机策略。在群体层面,在可实现性、精确优化、通用部署信息及分布保持部署的条件下,直接行为克隆与世界-动作模仿均能恢复观测行为策略。因此,未来预测仅改变学习因子分解方式,而不改变无约束外部策略类别或理想模仿目标。动作条件世界模型学习则通过预测指定动作下的结果,并基于控制目标对结果进行比较来实现。我们刻画了不依赖候选动作的未来模型的不可约动作特定预测误差,确定了世界-动作联合模型可恢复干预正向模型的条件,并指出观测演示通常无法识别动作效果。最后,我们构建了一个环境族,其中所有观测学习者均存在最坏情况遗憾值,而一次有信息的干预即可实现零遗憾值。因此,关键区别在于:为策略优化而预测与观测行为相关的未来,与预测指定动作的后果。

英文摘要

World-action models predict a future outcome and then infer an associated action. Although this factorization can improve representation learning and data efficiency, it is unclear whether it provides stronger control capability than direct behavior cloning when both are trained from the same observational demonstrations. We compare a direct behavior-cloning policy, an imitation-trained world-action policy, and a policy optimized with an action-conditioned world model. At the controller-class level, every world-action policy can be flattened into a direct stochastic policy with the same closed-loop trajectory distribution. At the population level, under realizability, exact optimization, common deployment information, and distribution-preserving deployment, direct behavior cloning and world-action imitation both recover the observational behavior policy. Thus, future prediction changes the learning factorization but not the unrestricted external policy class or ideal imitation target. Action-conditioned world-model learning differs by predicting outcomes under specified actions and comparing them through a control objective. We characterize the irreducible action-specific prediction error of future models that do not condition on the candidate action, identify conditions under which a world-action joint can recover an interventional forward model, and show that observational demonstrations do not identify action effects in general. Finally, we construct an environment family in which every observational learner has positive worst-case regret, whereas one informative intervention permits zero regret. The key distinction is therefore between predicting futures associated with observed behavior and predicting consequences of specified actions for policy optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑