机器人世界模型对动作书写方式并非不变
Robot World Models Are Not Invariant to How the Actions Are Written
浏览论文内容
中文总结 AI 辅助
机器人世界模型对动作参数化敏感,不同编码导致性能崩溃,提出测试识别问题并通过目标平均修复。
中文摘要 AI 辅助
机器人策略使用两种动作参数化之一进行训练:绝对关节目标,或相对于当前状态的增量。这一选择是机器人学习中的实时工程决策,而以动作为条件的世界模型会静默地继承这一选择。我们表明这种继承是灾难性的。在一个参数化上训练的潜在动力学模型,当接收到以另一种参数化书写的相同指令轨迹时,会崩溃:在三个机器人数据集和两种形态上,检索性能下降2.6-13.4倍,目标条件动作选择从53%降至15%,在PushT上,关于同一未来的两种信念近乎正交(cos=0.067,最差情况-0.377),因此预测器不会优雅地退化,而是回答了一个不同的问题。这不是通常意义上的分布偏移伪影:两种编码在给定联合输入时可相互重构,R^2=0.996,因此没有信息丢失,我们给出了区分有效重新参数化与有损摘要或传感器交换的测试。该测试拒绝了我们提出的四个轴中的三个。缺陷存在于动作通道,而视觉模型的不变性文献未检查该通道:那里的工作涉及裁剪、抖动和相机姿态,而命令的参数化未受审计。修复方法是对两种编码进行平均,而平均的位置很重要。对目标进行平均本身就能恢复任务性能;对输出进行平均,由于凹性对概率是安全的,但不适用于方向值预测,其中归一化均值可能低于轨道中的每个成员。目标平均遗留的是尾部:最坏情况一致性保持在0.78,不一致惩罚将其收窄至0.995,在潜在展开中,这是最坏情况侵蚀与保持之间的差异。在PushT上,仅平均无法修复该轴。
英文摘要
A robot policy is trained with one of two action parameterizations: absolute joint targets, or deltas relative to the current state. The choice is a live engineering decision in robot learning, and a world model conditioned on actions inherits it silently. We show the inheritance is catastrophic. A latent dynamics model trained on one parameterization and handed the identical commanded trajectory written in the other collapses: retrieval degrades by 2.6-13.4x across three robot datasets and two morphologies, goal-conditioned action selection falls from 53% to 15%, and on PushT the two beliefs about the same future are near-orthogonal (cos = 0.067, worst case -0.377), so the predictor does not degrade gracefully, it answers a different question. This is not a distribution-shift artifact in the usual sense: the two encodings are mutually reconstructible at R^2 = 0.996 given the joint input, so no information is lost, and we give the test that separates a valid re-parameterization from a lossy summary or a sensor swap. The test rejected three of the four axes we proposed. The defect lives in the action channel, which the invariance literature for visual models does not examine: work there concerns crops, jitter and camera pose, while the parameterization of the commands goes unaudited. The repair is averaging over the two encodings, and where it goes matters. Averaging the objective restores task performance by itself; averaging the outputs, safe for probabilities by concavity, is not available for direction-valued prediction, where the normalized mean can score below every member of the orbit. What objective-averaging leaves behind is the tail: worst-case agreement stays at 0.78, a disagreement penalty closes it to 0.995, and over a latent rollout it is the difference between a worst case that erodes and one that holds. On PushT, averaging alone does not repair the axis.
发表机构
- Hassana Labs(哈萨纳实验室)
- University of Oxford(牛津大学)
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。