When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
当模拟欺骗时:一个仿真到现实的基准和领域随机化的强化学习配方用于工具使用智能体
机构 * Arizona State University(亚利桑那州立大学) ; University of Southern California(南加州大学) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Pennsylvania(宾夕法尼亚大学) ; Adobe Research(Adobe研究)
AI总结 本文研究了工具使用POMDP中的仿真到现实差距,提出了一种领域随机化的强化学习方法,通过扰动增强轨迹提升鲁棒性,缩小了工具使用智能体在现实部署中的性能差距。
Comments Dataset, code, and benchmark leaderboard are available at https://github.com/WillChow66/robustbench-tc-release.git and https://huggingface.co/spaces/willchow66/robustbench-tc-leaderboard