DLAM:带时间约束的分布隐式动作模型
DLAM: Distributional Latent Actions with Temporal Constraints
浏览论文内容
中文总结 AI 辅助
该研究针对VLA模型数据稀缺问题,提出带时间约束的分布隐式动作模型DLAM,通过特定约束优化隐式动力学,提升了视频重建效果及机器人操作任务的策略性能。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型受限于稀缺的带动作标签的机器人数据,而无动作视频能提供丰富的物理变化观测。隐式动作模型可提取此类先验,但基于重建训练的编码可能在无机器人动作联合生成所需结构的情况下预测未来观测。现有结构化方法添加了时间约束,但保留了确定性转换点,因此局部推断转换中的残差误差会在递归组合下传播并累积。我们引入DLAM,一种分布隐式动作模型,将每个转换表示为对角高斯分布。以参考帧为条件的重建将均值锚定在观测到的视觉变化上,而对等间隔三元组的归一化组合与反转同时约束均值和维度方向的方差。方差组合使用轻量级共享相关系数来解释共享中间帧的相邻转换间的依赖关系,而反转操作会取均值的反数并保留方差。对于下游策略学习,我们冻结编码器并训练流匹配策略以联合生成均值转换序列和机器人动作。在保留的转换上,DLAM比现有隐式动作基线学习到更具时间一致性的隐式动力学,并在保留的视频上实现更强的直接和累积重建。在相同的受控π₀迁移协议下,它还提升了在MetaWorld MT50、LIBERO及真实世界操作任务上的策略性能。受控 ablation 实验表明,归一化均值约束贡献了大部分重建增益,而学习到的方差和感知相关性的组合则为下游控制提供了互补改进。
英文摘要
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions may propagate and compound under recursive composition. We introduce DLAM, a distributional latent-action model that represents each transition as a diagonal Gaussian. Reconstruction conditioned on the reference frame grounds the mean in observed visual change, while normalized composition and reversal over equal-gap triplets constrain both the mean and dimension-wise variance. Variance composition uses a lightweight shared-correlation coefficient to account for dependence between adjacent transitions that share an intermediate frame, whereas reversal negates the mean and preserves the variance. For downstream policy learning, we freeze the encoder and train a flow-matching policy to jointly generate mean transition sequences and robot actions. On held-out transitions, DLAM learns more temporally consistent latent dynamics than existing latent-action baselines and achieves stronger direct and cumulative reconstruction on held-out videos. Under the same controlled $π_0$ transfer protocol, it also improves policy performance on MetaWorld MT50, LIBERO, and real-world manipulation tasks. Controlled ablations show that normalized mean constraints account for most of the reconstruction gain, while learned variance and correlation-aware composition provide complementary improvements in downstream control.