发表机构
Soochow University; University of Electronic Science and Technology of China; Tongji University(苏州大学; 电子科技大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长时程机器人操作中动作分块策略在关键时间步预测误差主导失败的问题,提出不确定性引导的稀疏精化框架UGR,通过解耦不确定性分支仅校正最不确定的时间步,在RoboTwin五个任务中四项达到最优,较ACT提升达13%。
AI 中文摘要
学习基于分块的视觉运动策略以完成长时程机器人操作仍然具有挑战性。最近的动作分块方法通过预测时间上延展的动作序列展现了有前景的性能。然而,它们的失败往往由少数关键时间步上的预测误差主导,而非整个动作分块上的均匀差预测,这使得均匀精化效率低下且针对性不足。为解决这一瓶颈,我们提出了不确定性引导的精化(UGR),一种用于基于分块的视觉运动策略的稀疏精化框架。具体而言,UGR遵循从粗到精的设计:它首先预测完整的动作分块,从粗粒度隐藏状态估计每步的时间不确定性,并仅对由二元掩码选择的最不确定的时间步应用残差校正。不确定性分支与粗动作预测器解耦,使得精化增益能够清晰地归因于不确定性引导的校正,而非额外的预测器容量。在RoboTwin基准的五个双臂操作任务上的广泛实验表明,UGR在四个任务上取得了最佳成功率,相比ACT基线提升了最多13个百分点的绝对性能,并且在消融研究中优于全分块和位置无关的分块精化。
英文摘要
Learning chunk-based visuomotor policies for long-horizon robot manipulation remains challenging. Recent action-chunking methods have shown promising performance by predicting temporally extended action sequences. However, their failures are often dominated by prediction errors at a small number of critical timesteps rather than uniformly poor predictions across the entire action chunk, making uniform refinement inefficient and insufficiently targeted. To address this bottleneck, we propose Uncertainty-Guided Refinement (UGR), a sparse refinement framework for chunk-based visuomotor policies. Specifically, UGR follows a coarse-to-refine design: it first predicts a full action chunk, estimates per-step temporal uncertainty from the coarse hidden states, and applies residual correction only to the most uncertain timesteps selected by a binary mask. The uncertainty branch is decoupled from the coarse action predictor, enabling clean attribution of the refinement gains to uncertainty-guided correction rather than additional predictor capacity. Extensive experiments on five dual-arm manipulation tasks from the RoboTwin benchmark show that UGR achieves the best success rate on four tasks, improves over the ACT baseline by up to 13% absolute, and outperforms both full-chunk and position-agnostic block refinement in ablation studies.