arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可证明安全的仿真到现实迁移

Provably Safe Sim-to-Real Transfer

Tingting Ni, Maryam Kamgarpour

arXiv 2609.01418首次发表:更新:

发表机构

EPFL(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对安全仿真到现实迁移问题,在无奖励安全强化学习框架下设计高效算法,可减少现实交互并保证安全探索,还能计算接近最优的可行策略,刻画了仿真器的益处。

AI 中文摘要

为缓解现实世界强化学习(RL)的样本复杂度,常见做法是先在样本成本低廉的仿真器中训练策略,再将学习到的策略部署到现实世界,期望其能有效泛化。但这种直接的仿真到现实迁移无法保证成功:由于仿真到现实的失配,在仿真器中训练的策略在现实世界中可能次优。纠正这种失配需要从现实系统收集数据,但在机器人、医疗保健等诸多应用中,该数据收集过程本身受安全约束,由此产生了安全仿真到现实迁移问题:智能体如何利用不完善的仿真器,同时确保安全的现实世界数据收集,并为目标系统学习接近最优的可行策略?我们在无奖励安全RL框架内构建安全仿真到现实迁移问题以解决该问题,设计了一种计算高效的算法,该算法利用仿真器信息以可证明的方式减少现实世界交互,同时确保安全探索,并能为任何潜在奖励函数计算接近最优的可行策略。我们的现实世界样本复杂度边界从仿真到现实失配的角度刻画了使用仿真器的益处。

英文摘要

We address safe sim-to-real transfer, in which an agent leverages an imperfect simulator and limited real-world interaction while ensuring safety throughout data collection in the real system. This problem arises in applications such as robotics and healthcare: simulators provide cheap data, but sim-to-real mismatch makes direct transfer unreliable, and collecting real-world data to correct this mismatch must itself be safe. Moreover, deployment objectives may vary across tasks, making it costly to collect new data for each reward function. We therefore formulate safe sim-to-real transfer as a reward-free safe reinforcement learning (RL) problem, in which data are collected once and reused to plan for arbitrary reward functions. We develop a computationally efficient algorithm that identifies where the simulator and real dynamics differ, uses certified simulator transitions where they are reliable, and estimates mismatched transitions from safely collected data. With high probability, every policy deployed during learning is feasible, and the collected data support the computation of a feasible and near-optimal policy for any reward function. When the simulator is uninformative, our algorithm recovers online reward-free safe RL while improving the best-known sample complexity by a factor of \(\widetildeΘ(H/ξ^2)\), where \(ξ\) is the safety margin of a baseline policy. When the simulator is accurate on most transitions, this improvement grows to \(\widetildeΘ(H^2|\mc S||\mc A|/(ξ^2|\mc B|))\), where \(|\mc B|\) denotes the size of the sim-to-real mismatch region.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑