arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于 zonotope 的软 Actor-Critic 几何方法用于运动学习

A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion Learning

Panagiotis Roditis, Panagiotis P. Filntisis, Petros Maragos

arXiv 2610.12113首次发表:更新:

发表机构

Robotics Institute, Athena Research Center; HERON - Hellenic Robotics Center of Excellence; School of Electrical & Computer Engineering, NTUA(雅典娜研究中心机器人研究所; 希腊卓越机器人中心; 雅典国立技术大学电气与计算机工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出 GeZo-SAC 方法,通过 zonotope 几何表示调整 Critic 悲观程度,在 MuJoCo-v5 运动任务上取得高回报,且执行效率高、高估偏差小。

AI 中文摘要

离线策略 Actor-Critic 方法通过取两个 Critic 的最小值来控制高估偏差,该方法在所有场景中使用相同的聚合规则,而不考虑两个 Critic 的分歧程度。我们提出 GeZo-SAC,该方法使用辅助几何表示来使 Critic 的悲观程度适应状态和动作。除了标量值外,每个 Critic 还预测一组定义 zonotope 的生成器。沿采样方向探测该 zonotope 可得到几何宽度,将其作为悲观偏移量从每个 Critic 值中减去,同时得到两个 Critic 之间的分歧度量,该度量与对数求和指数聚合。该分歧控制两个 Critic 的组合方式,当分歧增大时,组合方式从宽度加权平均向常规最小值过渡。推理阶段,部署的策略是未修改的 SAC Actor,因为生成器仅在 Critic 侧使用。在四个 MuJoCo-v5 运动基准任务和六个离线策略基线方法中,GeZo-SAC 在 Ant-v5 和 Hopper-v5 上实现了最高平均回报,在其余任务上与其他方法相比仍具有竞争力。我们的分析进一步表明,GeZo-SAC 在所有评估方法中实现了每米最低的平均执行器功和动作 effort,同时在四个环境中保持接近零的测量高估频率。

英文摘要

Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics. This uses the same aggregation rule everywhere, regardless of how the critics disagree. We propose \textbf{GeZo-SAC}, which uses auxiliary geometric representations to adapt critic pessimism to the state and action. Alongside its scalar value, each critic predicts a set of generators defining a zonotope. Probing this zonotope along sampled directions provides a geometric width, "subtracted from each critic value as a pessimistic offset, and a measure of disagreement between the two critics, aggregated with log-sum-exp. This disagreement controls how the critics are combined, moving from a width-weighted average toward the usual minimum as disagreement increases. At inference, the deployed policy is an unmodified SAC actor, since the generators are used only on the critic side during training.Across four MuJoCo-v5 locomotion benchmarks and six off-policy baselines, GeZo-SAC achieves the highest mean return on Ant-v5 and Hopper-v5 and remains competitive with other methods on the remaining tasks. Our analysis further shows that GeZo-SAC achieves the lowest average actuator work and action effort per metre among the evaluated methods, while maintaining near-zero measured overestimation frequency across all four environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑