arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36250cs.LGcs.RO

动作分块近端策略优化与反馈校正

Action Chunking Proximal Policy Optimization with Feedback Correction

  • Cornell University(康奈尔大学)
  • Seoul National University(首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Sanghyun Hahn, Jonghyun Choi

中文总结 AI 辅助

提出ACPPO及ACPPO-Corr,通过分块actor与反馈校正器结合,在25个模拟机器人任务中实现最强综合性能,验证了动作分块在在线PPO中的有效性。

中文摘要 AI 辅助

动作分块通过在强化学习中选择短动作序列而非单个动作来提供时间抽象,但许多现有方法在高维机器人控制中面临两个局限性。首先,许多方法依赖于动作块上的值函数,随着动作维度和分块长度的增加,这些值函数可能难以学习。其次,开环执行分块会丢失块内反馈,限制了在接触密集任务中的反应能力。我们提出了动作分块PPO(ACPPO),这是一种PPO扩展,使用分块actor同时保留标准状态值critic,从而避免分块Q函数。我们进一步提出了ACPPO-Corr,它通过逐步反馈校正器增强分块规划器,在线调整每个分块内的计划动作。在来自IsaacGym和Bi-DexHands的25个模拟机器人任务中,涵盖运动、手臂操作和灵巧手-物体交互,ACPPO-Corr在评估方法中实现了最强的综合性能,并在决策频率敏感和决策频率中性的任务子集上均表现最佳。消融实验表明,适度的分块长度效果最佳,校正器正则化对于平衡分块级规划与局部反馈至关重要。这些结果表明,当分块级规划与闭环校正相结合时,动作分块可以在在线PPO中有效。代码可在以下网址获取:此https URL。

英文摘要

Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First, many rely on value functions over action chunks, which can be difficult to learn as action dimensionality and chunk length grow. Second, executing chunks open-loop removes within-chunk feedback, limiting reactivity in contact-rich tasks. We present Action Chunking PPO (ACPPO), a PPO extension that uses a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts planned actions online within each chunk. Across 25 simulated robotics tasks from IsaacGym and Bi-DexHands, spanning locomotion, arm manipulation, and dexterous hand-object interaction, ACPPO-Corr achieves the strongest aggregate performance among evaluated methods and performs best on both decision-frequency-sensitive and decision-frequency-neutral task subsets. Ablations show that moderate chunk lengths work best and that corrector regularization is important for balancing chunk-level planning with local feedback. These results suggest that action chunking can be effective in online PPO when chunk-level planning is paired with closed-loop correction. The code is available at: https://github.com/hshhahn/ACPPO.

补充信息

↑