发表机构
Zhejiang University; LimX Dynamics(浙江大学; LimX动力学公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究模拟与现实差距问题,提出“世界翻译”方法,利用模拟器和学习动力学互补优势,反向提取动力学信息并跨域转换,实验表明该方法能更准确建模,在特定情况下优势明显,还证实可改进策略转移。
AI 中文摘要
在现实世界中部署经过模拟训练的机器人策略时,模拟与现实之间的差距仍然是一个基本挑战。实到模拟方法从现实角度缩小了这一差距,从现实数据中学习过渡动力学以构建更逼真的数字世界,其中学习到的动力学模型是主要实例。然而,此类方法面临部分可观测性问题,相同观测可能因不可观测因素而导致不同过渡。现有方法假定这些因素可从观测历史中恢复,但在观测历史无信息时可能失败。为解决此限制,我们提出了“世界翻译”,利用模拟器和学习到的动力学的互补优势。我们从观测到的过渡中反向提取不可观测的动力学信息,然后将此特征作为非配对域翻译问题在模拟与现实之间进行转换,以在转移域风格的同时保留动力学内容。在人形机器人、四足机器人和操纵器平台上的实验表明,我们的方法比基线实现了更准确的动力学建模,在无法从观测历史中恢复不可观测因素时收益最大。在Go2四足机器人上的实际机器人部署证实了改进的策略转移。
英文摘要
The gap between simulation and reality remains a fundamental challenge in deploying simulation-trained robotic policies in the real world. Real-to-sim methods narrow this gap from the real side, learning transition dynamics from real data to build a more realistic digital world. Learned dynamics models are their dominant instance. Such methods, however, face a partial observability problem: the same observation may branch to different transitions due to unobservable factors. Existing methods assume these factors can be recovered from observation history. However, this may fail whenever observation history is uninformative, such as a sudden contact event with no prior warning. To address this limitation, we propose \textit{World Translation}, which exploits a complementary strength of simulators and learned dynamics. Simulators are deterministic but physically imperfect, while learned models are accurate but underdetermined under partial observability. Rather than predicting transitions forward from history, we extract the unobservable dynamics information backward from an observed transition, then translate this feature across simulation and reality as an unpaired domain-translation problem that preserves dynamics content while transferring domain style. Experiments across humanoid, quadruped, and manipulator platforms show that our method achieves more accurate dynamics modeling than baselines, with the largest gains when unobservable factors cannot be recovered from observation history. Real-robot deployment on Go2 quadruped confirms improved policy transfer.
Comments8 pages, 8 figures