TERRA:通过时间效果表示与关系对齐学习可迁移的潜在动作
TERRA: Learning Transportable Latent Actions through Temporal Effect Representation and Relational Alignment
- Institute of Artificial Intelligence, University of Central Florida(中佛罗里达大学人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TERRA通过时间效果表示与关系对齐学习可迁移的潜在动作,在LIBERO基准上达到93.4%的成功率,优于UniVLA的91.8%。
AI中文摘要:
潜在动作利用从视觉转换中推断出的类似动作的编码来监督机器人策略,其有效性取决于两个问题:编码从转换中保留了什么,以及当它在不同的初始状态下被重用时,是否仍然表示相同的意思。第一个问题是时间上的张力:终点差异忽略了运动如何展开,而完整序列则引入了干扰性变化。第二个问题则被重建方法所遗留,重建方法只观察到潜在编码与其来源状态的结合。我们认为,这两个问题可以在同一框架中得到解答。TERRA(时间效果表示与关系对齐)通过一个紧凑的时间效果来描述转换,该效果包括净特征变化和低阶窗口内动力学分量,并从这个效果中学习一个连续的潜在编码。相同的效果空间随后作为重用的参考:效果锚定迁移(EAT)在其他初始状态下解码潜在编码,并将产生的效果锚定到其源状态观察到的效果上,从而使潜在编码由其在跨上下文中的行为塑造,而不仅仅是由其来源的转换塑造。使用冻结的线性读取器,TERRA在动作预测上比UniVLA和LAPA风格的基线更准确,在视觉干扰下退化更慢,并且当接收上下文距离更远时,保持迁移的转换对捐赠动作的忠实度;一个相同预算的对照实验表明,这些收益主要来自EAT。在匹配的预训练规模下,完整系统在LIBERO上达到93.4%的平均成功率,而UniVLA为91.8%。
英文摘要:
Latent actions supervise robot policies with action-like codes inferred from visual transitions, and their usefulness hinges on two questions: what a code keeps from a transition, and whether it still means the same thing when reused in a different initial state. The first is a tension in time: an endpoint difference discards how motion unfolds, while the full sequence admits nuisance variation. The second is left open by reconstruction, which only ever observes a latent together with the state it came from. We argue that both questions can be answered in the same place. TERRA (Temporal Effect Representation and Relational Alignment) describes a transition by a compact temporal effect, its net feature change together with a low-order within-window dynamics component, and learns a continuous latent from this effect. The same effect space then serves as the reference for reuse: Effect-Anchored Transport (EAT) decodes a latent in other initial states and anchors the resulting effect to the one observed at its source, so that the latent is shaped by what it does across contexts rather than only by the transition it came from. With frozen linear readers, TERRA predicts actions more accurately than UniVLA and a LAPA-style baseline, degrades more slowly under visual distractors, and keeps transported transitions faithful to the donor action as the recipient context moves farther away; a same-budget control shows that these gains come largely from EAT. At matched pretraining scale, the complete system reaches 93.4% average success on LIBERO, compared with 91.8% for UniVLA.