arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07065cs.ROcs.AIcs.CVcs.HCcs.LG

AutoIntervene:用于动作分块模仿学习策略的校准干预方法

AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies

  • Australian Center For Robotics, The University of Sydney(悉尼大学澳大利亚机器人中心)
  • College of Connected Computing, Vanderbilt University(范德堡大学连接计算学院)

机构由 AI 辅助整理,请以论文原文为准。

Jinhe Tang, Weiming Zhi

AI总结:

AutoIntervene是一种在线框架,通过校准的双向切换阈值在动作分块模仿学习策略与操作员间切换控制,提升了真实世界双臂操作任务的适应后成功率并降低了操作员控制时间。

AI中文摘要:

动作分块视觉运动策略通过从演示中学习,预测短动作序列而非单步指令,以此提升时间一致性。然而感知误差与执行漂移会使机器人脱离演示分布,而策略仍会生成与观测状态不符的平滑动作分块。本文提出AutoIntervene,一种在部署时选择性在动作分块策略与操作员间切换控制的在线框架。AutoIntervene基于成功任务执行构建的视觉-动作支持记忆评估拟议分块,结合视觉相似度与拟议动作和参考动作的一致性。阶段内支持控制当前任务阶段内从策略到操作员的切换,全局支持控制操作员恢复后回归策略控制。该方法从保留的专家演示评估分数的经验分位数中为两个方向校准独立切换阈值,避免手动调整分数 cutoff。从成功 rollout 中保留的干预片段针对学习者诱导的状态,为后续策略更新提供纠正监督。在真实世界双臂操作任务上的实验表明,与手动干预相比,其适应后任务成功率更高,操作员控制时间更短。视频与额外结果可在该 https URL 获取。

英文摘要:

Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather than single-step commands. Yet perception errors and execution drift can move the robot outside the demonstration distribution, while the policy continues to produce smooth action chunks that are inconsistent with the observed state. We present AutoIntervene, an online framework that selectively transfers control between an action-chunking policy and an operator during deployment. AutoIntervene evaluates proposed chunks against a visual-action support memory built from successful task executions, combining visual similarity with consistency between proposed and reference actions. Phase-local support governs policy-to-operator transfer within the current task phase, whereas global support governs the return to policy control after operator recovery. We calibrate separate switching thresholds for the two directions from empirical quantiles of evaluation-level scores on held-out expert demonstrations, avoiding direct manual tuning of score cutoffs. Intervention segments retained from successful rollouts target learner-induced states and provide corrective supervision for subsequent policy updates. Experiments on real-world bimanual manipulation tasks show higher post-adaptation task success and lower operator-control time than manual intervention. Videos and additional results are available at https://aus.bot/research/autointervene/.

补充信息

↑