发表机构
Kyung Hee University(庆熙大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究重新审视时序正则化,证明其能提供空间平滑性,并提出CATS方法结合时序惩罚与线性渐进提升,在减少动作振荡的同时保持任务性能,经仿真和实物实验验证有效。
AI 中文摘要
深度强化学习策略可能产生不平滑的动作振荡,从而阻碍其在物理机器人上的部署。现有的基于架构和基于惩罚的方法通过直接降低对状态输入变化的敏感性来寻求空间平滑性,但其宽泛的约束在追求更强平滑性时可能会降低任务性能。时序正则化则沿观测到的转移约束动作差异,但被认为无法在观测噪声下提供所需的空间平滑性。我们通过证明时序惩罚限制了共享下一状态的当前状态之间的期望动作差异,从而重新审视了这一假设,揭示了一种经验上可扩展至空间平滑性的空间效应。基于这一发现,我们提出了仅使用时序平滑性的动作条件化(CATS),该方法将时序惩罚与线性渐进提升相结合。我们强调了时序正则化在提供空间平滑性的同时,比显式空间正则化更好地保持任务性能的能力。通过线性渐进提升,CATS允许策略在逐步平滑其动作之前学习有益行为,从而改善回报保持以及时序和空间平滑性。在仿真和现实世界的实验中表明,CATS在几乎不增加计算开销的情况下,显著减少了动作振荡,同时不降低任务性能。
英文摘要
Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad constraints can degrade task performance as stronger smoothing is pursued. Temporal regularization instead constrains action differences along observed transitions, but has been considered unable to provide the spatial smoothness needed under observation noise. We revisit this assumption by proving that the temporal penalty bounds the expected action differences between current states sharing a next state, revealing a spatial effect that empirically extends to spatial smoothness. Building on this finding, we propose Conditioning for Action using only Temporal Smoothness (CATS), which combines a temporal penalty with linear ramp-up. We highlight temporal regularization's ability to provide spatial smoothness while better preserving task performance than explicit spatial regularization. Through linear ramp-up, CATS allows the policy to learn rewarding behavior before progressively smoothing its actions, improving return preservation and both temporal and spatial smoothness. Experiments in both simulation and the real world show that CATS substantially reduces action oscillation without degrading task performance, with little computational overhead.
CommentsSubmitted to ICRA 2027