学习高风险高精度运动控制
Learning High-Risk High-Precision Motion Control
查看机构详情
- Aalto University(阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对高风险高精度运动控制问题,提出基于优势加权回归的SCOOT算法,通过精英样本优化、混合专家策略和距离正则化,在计算台球中实现高精度击球与多策略发现。
中文摘要 AI 辅助
深度强化学习(DRL)算法用于运动控制时,通常在顺序决策任务上进行评估和基准测试,在这些任务中,不精确的动作可以通过后续动作进行修正,从而允许在动作带有噪声的情况下仍能获得高回报。相比之下,我们关注一类研究不足的高风险、高精度运动控制问题,其中动作带有不可逆的后果,导致状态-动作奖励景观中出现尖锐的峰值和脊线。以计算台球作为此类问题的代表性示例,我们提出并评估了状态条件射击(SCOOT),这是一种新颖的DRL算法,基于优势加权回归(AWR)并进行了三项关键修改:1)仅使用精英样本进行策略优化,使策略能够更好地锁定稀有的高奖励动作样本;2)利用混合专家(MoE)策略,允许根据状态在奖励景观模式之间切换;3)添加距离正则化项和学习课程,以鼓励在适应最有利样本之前探索多样化的策略。我们展示了我们的特性在学习基于物理的台球击球中的性能,证明了高动作精度以及针对给定球配置发现多种击球策略的能力。
英文摘要
Deep reinforcement learning (DRL) algorithms for movement control are typically evaluated and benchmarked on sequential decision tasks where imprecise actions may be corrected with later actions, thus allowing high returns with noisy actions. In contrast, we focus on an under-researched class of high-risk, high-precision motion control problems where actions carry irreversible outcomes, driving sharp peaks and ridges to plague the state-action reward landscape. Using computational pool as a representative example of such problems, we propose and evaluate State-Conditioned Shooting (SCOOT), a novel DRL algorithm that builds on advantage-weighted regression (AWR) with three key modifications: 1) Performing policy optimization only using elite samples, allowing the policy to better latch on to the rare high-reward action samples; 2) Utilizing a mixture-of-experts (MoE) policy, to allow switching between reward landscape modes depending on the state; 3) Adding a distance regularization term and a learning curriculum to encourage exploring diverse strategies before adapting to the most advantageous samples. We showcase our features' performance in learning physically-based billiard shots demonstrating high action precision and discovering multiple shot strategies for a given ball configuration.