PAC-ACT:用于动作分块变换器的训练后演员评论家方法
PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers
浏览论文内容
中文总结 AI 辅助
针对精密工业接触操纵问题,提出PAC-ACT框架,通过在块级别重新制定策略优化、构建特定架构并引入混合行为先验约束,提升了机器人策略在姿态扰动和接触力约束下的性能,实验验证了其有效性。
中文摘要 AI 辅助
精密工业接触操纵需要在姿态扰动和接触力约束下有可靠的机器人策略。视觉-语言-动作模型具有广泛的通用性,但通常会带来高推理延迟和高GPU内存成本,而视觉动作分块策略更适合实时工业控制。然而,这些策略通常通过行为克隆进行训练,在富含接触的任务中会受到分布偏移的影响。本文提出了PAC-ACT,一种用于预训练动作分块变换器策略的强化学习训练后框架。PAC-ACT在块级别重新制定策略优化,构建了一个ACT转移的演员评论家架构,并引入了混合行为先验约束,以在在线微调期间保留预训练的动作分布。在工业精密接触基准上的实验表明,PAC-ACT提高了任务成功率、接触稳定性和力安全性,同时保持了低延迟和低GPU内存使用。在轮廓任务上,PAC-ACT显著降低了峰值接触力,并将60N以上力读数的比例降低了46倍。稀疏奖励消融进一步表明,所提出的行为先验约束能够在随机初始姿态下进行有效探索。
英文摘要
Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization but often introduce high inference latency and GPU-memory cost, while vision-action chunking policies are more suitable for real-time industrial control. However, these policies are usually trained by behavior cloning and suffer from distribution shift in contact-rich tasks. This paper proposes PAC-ACT, a reinforcement-learning post-training framework for pretrained Action Chunking Transformer policies. PAC-ACT reformulates policy optimization at the chunk level, constructs an ACT-transferred actor-critic architecture, and introduces a hybrid behavior-prior constraint to preserve the pretrained action distribution during online fine-tuning. Experiments on industrial precision-contact benchmarks show that PAC-ACT improves task success, contact stability, and force safety while retaining low latency and low GPU-memory usage. On the Contour task, PAC-ACT significantly reduces peak contact force and decreases the proportion of force readings above 60 N by 46 times. Sparse-reward ablations further show that the proposed behavior-prior constraint enables effective exploration under randomized initial poses.
发表机构
- LeRobot
机构由 AI 辅助整理,请以论文原文为准。