arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

识别机器人世界模型中的习惯、物理与干扰

Identifying Habit, Physics, and Nuisance in Robot World Models

Jinting Hang, Zhenhui Cai

arXiv 2609.09210首次发表:更新:

发表机构

Harvest Praxis(Harvest Praxis)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对遥操作演示多模态问题,提出结构因果模型分解习惯、物理与干扰,通过冻结物理层更新薄接口的适应规则,在多个数据集上提升低样本迁移性能。

AI 中文摘要

遥操作演示通常是多模态的,即使在给定执行动作的情况下底层动力学几乎是确定性的。我们认为,这种多模态通常混合了三个因素——操作者在动作选择中的习惯、共享的物理规律以及观测干扰——而纠缠的下一观测预测器会吸收这三者。我们通过结构因果模型 a=g(h,z,u), z'=f(z,a), o=r(z,c) 形式化了这一分解,并用互补干预进行测试:在固定状态下替换或打乱动作会显著增加下一状态误差,而外观和相机变化则不应如此;习惯感知的逆向评分在不重写动力学的情况下提高了可行过去轨迹的排序。相关的适应规则是冻结共享的物理读出层,仅更新一个薄接口。在 StackCube、DROID 和 RH20T 上,该规则相比从头训练改善了低样本迁移,在受污染的适应数据下保持了更干净的动力学,并通过多视角和多步检查从本体感觉扩展到像素观测。我们不将潜在动作等同于操作者习惯,也不针对大规模视频生成基准。

英文摘要

Teleoperated demonstrations are often multimodal even when the underlying dynamics are nearly deterministic given the executed action. We argue that this multimodality typically mixes three factors--operator habit in action selection, shared physics, and observation nuisance--and that entangled next-observation predictors absorb all three. We formalize the split with a structural causal model a=g(h,z,u), z'=f(z,a), o=r(z,c), and test it with complementary interventions: replacing or shuffling actions at fixed state sharply increases next-state error, whereas appearance and camera changes should not; habit-aware reverse scoring improves ranking of feasible pasts without rewriting the dynamics. The associated adaptation rule is to freeze a shared physics readout and update only a thin interface. On StackCube, DROID, and RH20T this rule improves low-shot transfer relative to training from scratch, retains cleaner dynamics under corrupted adaptation data, and extends from proprioception to pixel observations with multi-view and multi-step checks. We do not equate latent actions with operator habit, and we do not target large-scale video generation benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑