发表机构
Beijing University of Posts and Telecommunications; Institute of Automation, Chinese Academy of Sciences; Shenzhen University; Tsinghua University(北京邮电大学; 中国科学院自动化研究所; 深圳大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SPRII利用交互间关系作为弱监督,塑造持久上下文表征,在多个任务上平均提升下游性能超10%,持久属性读出超15%。
AI 中文摘要
世界模型从交互经验中学习环境动态。这些动态依赖于当前状态和动作,以及跨交互持续存在的属性。然而,标准预测训练可能仅利用局部证据来减少误差,而不将持久信息组织为可复用的上下文。我们提出SPRII,一种训练原则,利用交互之间的关系作为持久上下文的弱监督,同时保留学习者的原生目标。例如,同一系统的不同轨迹共享持久属性,即使其状态和动作不同。SPRII利用这种关系指导上下文学习,无需数值属性标签。两个可组合组件鼓励来自相关交互的上下文保持一致(Align),并使用一个交互的上下文来预测另一个交互的未来(Cross)。我们的分析区分了三个相互关联的问题:在学习的上下文中可访问哪些持久信息(Formation)、该上下文如何影响固定预测器(Use),以及它是否减少任务误差(Value)。一个阶段的成功并不能保证下一阶段的成功。受控实验表明,更可靠的关系改善表征组织,但添加共享属性约束可能减少对仍共享属性的访问。上下文替换在固定模型权重下改变预测,而历史收益取决于预测范围和读出方式。评估涵盖十三个设置,包括受控物理系统、公共动态任务、机器人和触觉数据,以及伙伴交互,跨越多个学习者家族。相对于相应基线,SPRII在下游任务性能上平均提升超过10%,在持久属性读出上平均提升超过15%。项目页面可在该https URL获取。
英文摘要
World models learn environment dynamics from interaction experience. These dynamics depend on the current state and actions, as well as on properties that persist across interactions. Yet standard predictive training can reduce error using local evidence alone, without organizing persistent information into reusable context. We introduce SPRII, a training principle that uses relations between interactions as weak supervision for persistent context while retaining the learner's native objective. For example, different trajectories of the same system share persistent properties even when their states and actions differ. SPRII uses such relations to guide context learning without numerical property labels. Two composable components encourage contexts from related interactions to agree (Align) and use one interaction's context to predict another's future (Cross). Our analysis distinguishes three linked questions: what persistent information is accessible in the learned context (Formation), how that context influences a fixed predictor (Use), and whether it reduces task error (Value). Success at one stage does not guarantee success at the next. Controlled experiments show that more reliable relations improve representation organization, but adding a shared-property constraint can reduce access to a property that remains shared. Context substitutions change predictions at fixed model weights, while the benefit from history depends on prediction horizon and readout. Evaluations span thirteen settings, including controlled physical systems, public dynamics tasks, robotic and tactile data, and partner interaction, across multiple learner families. Relative to the corresponding baselines, SPRII yields average gains of over 10% in downstream task performance and over 15% in persistent-property readout. The project page is available at https://persistent-learning-review.netlify.app/interactive.html.