重建并非行动:以行动为中心的潜在动力学建模
Reconstructing Is Not Acting: Action-Centric Latent Dynamics Modeling
浏览论文内容
中文总结 AI 辅助
针对潜在行动模型重建误差与行动质量不匹配的问题,提出以行动为中心的ACT-LAM框架,通过行动查询IDM和行动令牌FDM强化行动提取与利用,在VP$^2$基准上以更低开销提升成功率7.6%。
中文摘要 AI 辅助
潜在行动模型(LAMs)通过从视觉转换中推断潜在行动并重建未来状态,从无标签视频中学习行动表示。然而,我们发现了一个根本性的“重建-行动不匹配”问题:较低的重建误差并不必然带来更好的潜在动力学或下游性能。我们将这一不匹配归因于基于重建的潜在动力学建模中两个约束不足的方面:(i)逆动力学模型(IDM)未被明确鼓励区分与行动相关的转换与干扰性外观;(ii)正动力学模型(FDM)可能通过利用当前状态的预测捷径而未能充分利用推断出的潜在行动。为解决这两个局限,我们提出了ACT-LAM,一个轻量级的以行动为中心的框架,该框架同时加强了行动提取和行动利用。具体而言,其行动查询IDM(AQ-IDM)采用可学习的行动查询和门控聚合,以在无强信息瓶颈的情况下选择性地提取丰富的行动相关转换线索。其行动令牌FDM(AT-FDM)将潜在行动投影为行动令牌,这些令牌逐步与演化的状态表示交互,实现连续的状态感知行动条件化。ACT-LAM进一步简化特征处理,将模型能力集中于潜在动力学建模。在多个机器人数据集和VP$^2$基准上的大量实验表明,ACT-LAM在更少的可训练参数和更低计算开销下,实现了更强的潜在行动一致性、正动力学和下游视觉规划性能。特别是,ACT-LAM在聚合的VP$^2$成功率上超越了先前最先进的方法,提升了7.6%。代码可在https://this https URL获取。
英文摘要
Latent action models (LAMs) learn action representations from unlabeled videos by inferring latent actions from visual transitions and reconstructing future states. However, we identify a fundamental $\textbf{reconstruction-action mismatch}$: lower reconstruction error does not necessarily yield better latent dynamics or downstream performance. We attribute this mismatch to two underconstrained aspects of reconstruction-based latent dynamics modeling: (i) the inverse dynamics model (IDM) is not explicitly encouraged to distinguish action-related transitions from nuisance appearance, and (ii) the forward dynamics model (FDM) can underutilize the inferred latent action by exploiting predictive shortcuts from the current state. To address both limitations, we propose $\textbf{ACT-LAM}$, a lightweight action-centric framework that strengthens both action extraction and action utilization. Specifically, its Action Query IDM (AQ-IDM) employs learnable action queries and gated aggregation to selectively extract rich action-related transition cues without strong information bottlenecks. And its Action Token FDM (AT-FDM) projects latent actions into action tokens that progressively interact with evolving state representations, enabling continuous state-aware action conditioning. ACT-LAM further streamlines feature processing to concentrate model capacity on latent dynamics modeling. Extensive experiments on several robotic datasets and the VP$^2$ benchmark demonstrate stronger latent action consistency, forward dynamics, and downstream visual planning performance with fewer trainable parameters and lower computational overhead. In particular, ACT-LAM surpasses the previous state of the art by $\textbf{7.6%}$ on the aggregated VP$^2$ success rate. Codes at $\href{https://github.com/DingjieFu/ACT-LAM}{url}$.
发表机构
- Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。