发表机构
The University of Sydney; University of Oxford; The University of Western Australia(悉尼大学; 牛津大学; 西澳大利亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AIoT系统中强化学习模拟到现实的差距问题,开发了成本低于400美元的真实世界平台,通过硬件模拟键盘让边缘设备智能体玩游戏训练,实验显示模拟训练智能体部署后性能大幅下降,直接现实训练有可行性,为强化学习评估提供了基础。
AI 中文摘要
强化学习常用于提升包括自主物联网(AIoT)在内的自治系统性能。但在现实环境中进行强化学习代价高昂且有风险,多数研究在模拟环境中开展,这带来模拟到现实的可迁移性挑战。评估算法鲁棒性和差距是提升现实中强化学习性能的关键前提,机器人等行业已开发相关平台,但AIoT领域尚无通用基准平台。我们为此开发了一个用于研究AIoT中强化学习的真实世界平台,该平台使用成本低于400美元的商用组件和两台计算机,边缘设备上的智能体通过硬件模拟键盘在主机上玩视频游戏,以视觉输入为引导。实验结果表明模拟训练的智能体在实际部署后性能下降1160%。直接在现实世界用深度Q网络算法训练1000万步后可达到人类水平性能的约38%。该平台为现实AIoT系统中强化学习的定性和定量评估提供了基础。
英文摘要
Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT). However, the trial-and-error nature of RL, when conducted in real-world environments, is costly and hazardous in some scenarios. Consequently, the majority of RL research is conducted in simulation. This reliance introduces challenges related to the Sim-to-Real transferability. Evaluating the Sim-to-Real algorithmic robustness and the Sim-to-Real gap is a critical prerequisite for research aimed at improving RL performance in the real world. Therefore, industries such as robotics have developed concurrent simulation and physical platforms to facilitate this research. However, a universal Sim-to-Real benchmark platform for AIoT does not currently exist. To address these concerns, we developed a real-world AIoT platform for studying RL in AIoT. On this platform, an agent deployed on an edge device plays video games on a separate host computer via a hardware-emulated keyboard, guided by vision input. This platform uses commercially available components costing less than USD 400, together with two computers. Because the system's objective is game score maximization, it inherently mitigates safety risks associated with real-world RL deployments. Experimental results show the simulation-trained agent suffers a 1160% performance degradation relative to the human-level performance after real-world deployment, indicating a significant Sim-to-Real gap. Direct real-world training using the deep Q-network (DQN) algorithm achieves approximately 49% of human-level performance after 10 million training steps, demonstrating the feasibility of RL under real-world conditions. These results suggest that the proposed Sim-to-Real benchmark platform provides a substantial foundation for qualitative and quantitative evaluations of RL in real-world AIoT systems.
CommentsThe code for this paper is available upon the request from the author