arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

混合残差强化学习用于接触丰富的机器人书籍插入

Hybrid Residual Reinforcement Learning for Contact-Rich Robotic Book Insertion

Tianyuan Liu, Rutherford Agbeshi Patamia, Benjamin Champion, Akansel Cosgun, Richard Dazeley

arXiv 2609.19962首次发表:更新:

发表机构

Deakin University(迪肯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对紧密书架书籍插入的接触丰富控制难题,提出保留名义控制器并叠加残差PPO的混合方法,模拟成功率98.5%,实物成功率从26.7%提至63.3%,验证了几何与学习结合的有效性。

AI 中文摘要

将抓取的书籍放入紧密的书架是一个紧凑但困难的接触丰富控制问题:毫米级的位姿误差可以将几何上有效的接近转变为卡住、释放失败或未完全就位。我们研究抓取获取和全局接近之后的最终阶段,并探讨控制权应如何在已知几何和习得行为之间分配。我们的方法保留了一个名义任务空间控制器用于结构化插入和就位,而残差PPO提供有界局部修正并决定何时释放。只有短暂的打开-退回-重新关闭过渡是脚本化的。对于硬件上使用的最终策略,在512个固定条件下的部署匹配模拟评估中,三次独立训练运行的平均成功率为98.50%(样本标准差0.23个百分点),而名义控制为37.89%。在物理xArm7上,30个匹配条件下的60次试验显示了相同的定性优势:残差控制将成功率从26.7%提高到63.3%,失败次数从22次减少到11次,并在两个控制器不同的15个匹配条件中赢得了13个。鲁棒性测试表明,在初始化扰动高达1.5倍时,性能仍保持在87%以上,而非常紧密的间隙暴露了局部修正的几何极限。这些结果支持一种混合设计,其中几何保留了可靠的任务结构,学习集中在固定规则处理不佳的接触敏感行为上。

英文摘要

Placing a grasped book into a tight shelf is a compact but difficult contact-rich control problem: millimetre-scale pose error can turn a geometrically valid approach into jamming, failed release, or incomplete seating. We study this final phase after grasp acquisition and global approach, and ask how control authority should be divided between known geometry and learned behaviour. Our method retains a nominal task-space controller for structured insertion and seating, while residual PPO supplies bounded local corrections and decides when to release. Only the brief open-retreat-reclose transition is scripted. For the final policy used on hardware, a deployment-matched simulation evaluation over 512 fixed conditions yields 98.50 percent mean success (0.23 percentage-point sample SD) across three independent training runs, compared with 37.89 percent for nominal control. On the physical xArm7, 60 trials over 30 matched conditions show the same qualitative advantage: residual control raises success from 26.7 percent to 63.3 percent, reduces failures from 22 to 11, and wins 13 of the 15 matched conditions in which the two controllers differ. Robustness tests show that performance remains above 87 percent under initialization perturbations up to 1.5x, while very tight clearances expose the geometric limit of local correction. These results support a hybrid design in which geometry preserves reliable task structure and learning is concentrated on the contact-sensitive behaviour that fixed rules handle poorly.

Comments8 pages, 8figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑