DAMPER:面向平滑策略的回报优先梯度控制
DAMPER: Return-Prioritized Gradient Control for Smooth Policies
浏览论文内容
中文总结 AI 辅助
DAMPER通过冲突条件投影和自适应幅度控制,结合原生演员梯度与时间一致性梯度,减少连续控制中的动作振荡,在12个任务-骨干组合中均降低振荡并保持回报权衡。
中文摘要 AI 辅助
演员-评论家方法在连续控制中取得了强劲性能,但其策略可能产生高度振荡的动作。常见的补救措施是添加辅助平滑损失。然而,当这些损失的梯度相对于原生演员梯度较小时,其贡献可能可以忽略不计。此外,现有方法通常组合多个辅助损失,使损失平衡复杂化,且不一定改善回报-平滑度权衡。我们提出DAMPER(具有显式回报优先级的方向感知幅度受控投影),通过冲突条件投影和自适应幅度控制,将原生演员梯度与时间一致性梯度相结合。它移除与演员梯度相对抗的辅助分量,并相对于演员梯度范数缩放保留的时间方向,从而保持与原生演员梯度的正对齐。在六个连续控制任务上使用TD3和SAC进行的实验表明,在所有12个任务-骨干组合中,相对于原生智能体,动作振荡均有所减少,并在其中八个组合中取得了所比较方法中的最佳振荡得分,同时存在任务相关的回报权衡。
英文摘要
Actor-critic methods achieve strong performance in continuous control, but their policies can produce highly oscillatory actions. A common remedy is to add auxiliary smoothness losses. However, their contribution can be negligible when their gradients are small relative to the native actor gradient. Moreover, existing methods often combine multiple auxiliary losses, complicating loss balancing without necessarily improving the return-smoothness trade-off. We introduce DAMPER (Direction-Aware Magnitude-Controlled Projection with Explicit Return Priority), which combines the native actor gradient with a temporal-consistency gradient through conflict-conditioned projection and adaptive magnitude control. It removes the auxiliary component opposing the actor gradient and scales the retained temporal direction relative to the actor gradient norm, preserving positive alignment with the native actor gradient. Experiments with TD3 and SAC on six continuous-control tasks show reduced action oscillation relative to the native agents in all 12 task-backbone pairs and the best oscillation score among the compared methods in eight, with task-dependent return trade-offs.
发表机构
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。