面向离线视觉控制的定向时序表示
Directed Temporal Representations for Offline Visual Control
- University of Sheffield(谢菲尔德大学)
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai University of Finance and Economics(上海财经大学)
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DTRC利用定向时序拟度量从离线视觉轨迹学习控制表示,校准时间距离并利用进展信号训练目标条件策略,在多个视觉任务中超越基线。
AI中文摘要:
预测性世界模型为控制提供了紧凑的视觉表示。控制需要与时间可达性对齐的潜在几何结构,而不仅仅是预测相似性。我们引入了用于控制的定向时序表示(DTRC),该方法在冻结的LeWorldModel(LeWM)特征之上,从离线视觉轨迹中学习这种几何结构。DTRC在学习的控制表示上构建了一个定向时序拟度量。短程时间偏移校准距离尺度。自举目标将时间可达性扩展到更长的时间范围。动作条件一致性使表示与局部转移动态对齐。由此产生的距离估计了时间到达成本,其跨转移的变化定义了目标相对的时间进展。我们使用这一进展信号作为时间评论员,用于直接的目标条件策略学习。模型辅助目标在行为支持和动态一致性约束下提供了额外的训练时细化。在十个视觉控制任务中,DTRC相对于规划和直接策略基线实现了强大的目标条件控制性能。在四个LeWM任务上的保留诊断显示出一致的短程时间校准、任务相关的长程和方向结构,以及正面的转移级进展。时间监督在所有四个LeWM任务上改善了相同的流策略参数化,而所得策略在测试时无需迭代轨迹搜索即可直接行动。
英文摘要:
Predictive world models provide compact visual representations for control. Control requires a latent geometry aligned with temporal reachability rather than predictive similarity alone. We introduce Directed Temporal Representations for Control (DTRC), which learns such a geometry from offline visual trajectories on top of frozen LeWorldModel (LeWM) features. DTRC constructs a directed temporal quasimetric over the learned control representation. Short-range temporal offsets calibrate the distance scale. Bootstrapped targets extend temporal reachability across longer horizons. Action-conditioned consistency aligns the representation with local transition dynamics. The resulting distance estimates temporal reaching cost, and its change across a transition defines goal-relative temporal progress. We use this progress signal as a temporal critic for direct goal-conditioned policy learning. Model-assisted targets provide an additional training-time refinement under behavior-support and dynamics-agreement constraints. Across ten visual control tasks, DTRC achieves strong goal-conditioned control performance relative to planning and direct-policy baselines. Held-out diagnostics on the four LeWM tasks show consistent short-range temporal calibration, task-dependent long-range and directional structure, and positive transition-level progress. Temporal supervision improves the same flow-policy parameterization across all four LeWM tasks, while the resulting policy acts directly without iterative trajectory search at test time.