arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

模仿学习是否保留灵巧操作中的时间鲁棒性?不同任务执行速度下的专家-学习者对比

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds

Clinton Enwerem, John S. Baras, Calin Belta

arXiv 2609.01453首次发表:更新:

发表机构

Institute for Systems Research; University of Maryland(系统研究所; 马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比灵巧操作中模仿学习的ACT策略与脚本化专家,发现两者额定速度下成功率均为100%,但加速时ACT性能下降更显著,说明模仿学习未保留专家的时间鲁棒性。

AI 中文摘要

通过模仿学习得到的灵巧操作策略,通常会评估其对场景、物体或指令变化的鲁棒性,但较少考察其在不同任务执行速度下的性能。这使得学习者相对于其模仿的专家保留了多少时间鲁棒性尚不明确。我们在相同任务条件、初始条件采样和加速因子下对比专家与学习者,将评估实例化为ParcelStow任务,这是一个接触丰富的任务,机器人需抓取、重新定向并插入包裹。演示覆盖了抓取后操作阶段的加速范围。脚本化专家和从专家演示训练得到的ACT(Action Chunking with Transformers)策略,在额定速度下均实现100%的任务成功率。在演示范围内,它们的成功率出现分化:在最大速度下,专家成功率为84%,ACT成功率为53%。两个具有不同参数初始化的ACT策略表现出相似的性能下降,从额定速度到最大演示速度分别下降34和48个百分点,而专家仅下降16个百分点。阶段级分析显示,在最大演示速度下,ACT的47次失败中有35次是插入错位。在相对运动交接条件下,所有ACT抓取都能在自由空间中完成重新定向和转移时保持包裹,但仅64%完成整体任务,而专家抓取后完成率为95%。在所有评估的策略和速度下,414次无力闭合的抓取均未完成任务。因此,相同的额定任务成功率并不意味着在不同执行速度下保留了专家性能。代码、数据和评估脚本可在该httpsURL获取。

英文摘要

Dexterous manipulation policies learned by imitation are typically evaluated for robustness to variation in scenes, objects, or instructions, but their performance across task execution speeds is less often examined. This leaves open how much temporal robustness a learner retains relative to the expert it imitates. We compare an expert and learner under the same task conditions, initial-condition draws, and speedup factors. We instantiate the evaluation in ParcelStow, a contact-rich task in which the robot acquires, reorients, and inserts a parcel. The demonstrations span the speedup range for the manipulation phases after parcel acquisition. A scripted expert and an Action Chunking with Transformers (ACT) policy trained from the expert's demonstrations both achieve 100 percent task success at nominal speed. Their success rates diverge within the demonstrated range: at its maximum, expert success is 84 percent and ACT success is 53 percent. Two ACT policies with different parameter initializations show similar degradation, decreasing by 34 and 48 percentage points from nominal speed to the maximum demonstrated speed, compared with 16 points for the expert. Stage-level analysis shows that 35 of ACT's 47 failures at the maximum demonstrated speed are insertion misalignments. Under the relative-motion handoff, every ACT acquisition retains the parcel through reorientation and transfer in free space, but only 64 percent complete the overall task, compared with 95 percent after expert acquisition. Across all evaluated policies and speeds, none of the 414 acquisitions without force closure completes the task. Equal nominal task success therefore does not imply preservation of expert performance across execution speeds. Code, data, and evaluation scripts are available at https://github.com/coenwerem/parcelstow.

Comments19 pages, 10 figures, and 9 tables. Code, data, and evaluation scripts are available at https://github.com/coenwerem/parcelstow

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑