arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

透过位移坐标系:视觉-力觉精密装配的特权噪声蒸馏

Seeing Through the Displaced Frame: Privileged Noise Distillation for Vision-Force Precision Assembly

Ching-Hsiang Chang, Tzu-Yu Chuang, Yi-Hsiu Lee, Yi-Ting Chen, Yuan-Fu Yang, Min Sun

arXiv 2610.07745首次发表:更新:

发表机构

National Tsing Hua University; National Yang Ming Chiao Tung University(国立清华大学; 国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对精密装配中位姿噪声导致的坐标系偏移问题,提出利用特权教师和闭式重标注提供偏移监督,经行为克隆和DAgger训练,显著提升学生模型在噪声下的成功率。

AI 中文摘要

精密装配中的位姿误差不仅会破坏机器人所观测到的信息,还会破坏其执行动作所依据的坐标系。在FORGE基准测试中,官方基于状态的政策在真实位姿下成功率达到97%至99%,但在基准的σ=5毫米位姿噪声设置下,成功率仅为32%至60%。相同的估计位姿既进入观测又锚定动作坐标系,使得在接触前仅凭本体感觉状态无法识别该偏移。我们在训练期间通过两种方式提供这一缺失信息。一个特权教师模型在仿真中观测偏移,而干净示范则可以通过闭式解重新标注到位移坐标系中。部署的学生模型通过行为克隆后接一轮DAgger进行训练,测试时仅接收含噪声状态、原始力矩窗口和两个RGB相机。在未修改的FORGE任务上,教师路径学生模型在σ=0至5毫米范围内保持92%至99%的成功率,而六个非特权基线则降至2%至80%。一个匹配的行为克隆实验隔离了这种鲁棒性的来源。在相同的学生架构、数据预算和训练流程下,无偏移访问生成的示范在σ=5毫米时仅达到31.5%的成功率,而特权示范和重新标注示范分别达到88.7%和95.8%。在Franka上零样本部署时,学生模型在σ=5毫米时达到83.3%的汇总成功率,而最强基于状态的政策仅为34.4%。仅靠可部署的传感是不够的。鲁棒性需要编码对潜在坐标系补偿的监督。项目页面:此URL页面:此URL

英文摘要

Pose error in precision assembly can corrupt not only what a robot observes but also the coordinate frame in which it acts. On the FORGE benchmark, the official state-based policy succeeds in 97% to 99% of episodes with the true pose but only 32% to 60% at the benchmark's $σ=5$ mm pose-noise setting. The same estimated pose enters the observation and anchors the action frame, making the offset unidentifiable from proprioceptive state alone before contact. We supply this missing information during training in two ways. A privileged teacher observes the offset in simulation, while clean demonstrations can instead be relabelled into the displaced frame in closed form. The deployed student is trained with behaviour cloning followed by one DAgger round and receives only noisy state, a raw wrench window, and two RGB cameras at test time. On the unmodified FORGE tasks, the teacher-route student maintains 92% to 99% success across $σ=0$ to 5 mm, while six non-privileged baselines fall to 2% to 80%. A matched behaviour-cloning experiment isolates the source of this robustness. With the same student architecture, data budget, and training procedure, demonstrations generated without offset access yield only 31.5% success at $σ=5$ mm, whereas privileged and relabelled demonstrations reach 88.7% and 95.8%. Deployed zero-shot on a Franka, the student reaches 83.3% pooled success at $σ=5$ mm against 34.4% for the strongest state-based policy. Deployable sensing alone is insufficient. Robustness requires supervision that encodes compensation for the latent frame offset.Project page: https://drychang.github.io/displaced-frame/

Comments8 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑