arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28878cs.RO

在线仿真到现实适应:通过闭环系统建模

Online Sim-to-Real Adaptation via Closed-Loop System Modeling

  • Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuhao Huang, Samuel A. Moore, Boyuan Chen

AI总结:

OSRAM框架通过闭环系统建模在线调整参考指令,实现仿真到现实的适应,无需修改控制策略,有效减少残余跟踪误差并提升预测精度。

AI中文摘要:

仿真到现实的迁移已取得实质性进展,但仍可能产生在硬件上保持稳定和功能正常,却因残余动力学不匹配而导致跟踪精度下降的控制器。纠正这些误差通常需要识别底层系统动力学、调整控制策略,或返回仿真进行额外训练和微调,所有这些都可能需要大量数据和计算。我们提出OSRAM(在线仿真到现实适应:通过闭环系统建模),这是一个框架,它转而调整提供给现有控制器的参考指令。OSRAM将部署的机器人及其策略视为一个统一的闭环动力学系统,并直接从跟踪观测中学习其任务级指令响应行为。一个闭环动力学模型在仿真中的随机动力学上进行元训练,并在部署后利用有限的现实交互快速微调。然后,调整后的模型用于优化未来的参考指令,同时保持底层控制策略不变。我们在仿真和硬件上对双足速度跟踪和移动操作评估了OSRAM。结果表明,闭环建模在未见动力学下提高了预测和跟踪精度,而在线参考调整减少了不同控制目标和硬件配置下的残余仿真到现实跟踪误差。这些结果表明,调整机器人-策略闭环的行为为仿真到现实迁移提供了一种实用的替代方案,无需微调策略或识别完整物理动力学。更多信息可在此http URL找到。

英文摘要:

Sim-to-real transfer has made substantial progress, but can still produce controllers that remain stable and functional on hardware while suffering from degraded tracking accuracy due to residual dynamics mismatch. Correcting these errors typically requires identifying the underlying system dynamics, adapting the control policy, or returning to simulation for additional training and finetuning, all of which can require substantial data and computation. We propose OSRAM (Online Sim-to-Real Adaptation via Closed-Loop System Modeling), a framework that instead adapts the reference commands provided to an existing controller. OSRAM treats the deployed robot and its policy as a unified closed-loop dynamical system and learns its task-level command-response behavior directly from tracking observations. A closed-loop dynamics model is meta-trained across randomized dynamics in simulation and rapidly finetuned after deployment using limited real-world interaction. The adapted model is then used to optimize future reference commands while leaving the underlying control policy unchanged. We evaluate OSRAM on bipedal velocity tracking and loco-manipulation in simulation and on hardware. Results show that closed-loop modeling improves prediction and tracking accuracy under unseen dynamics, while online reference adaptation reduces residual sim-to-real tracking errors across different control objectives and hardware configurations. These results demonstrate that adapting the behavior of the robot-policy closed loop provides a practical alternative to finetuning the policy or identifying the full physical dynamics for sim-to-real transfer. More information can be found at http://generalroboticslab.com/OSRAM.

↑