arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25887cs.RO

什么是更好的课程:控制器塑造的抓取行为用于接触力敏感操作

What is the Better Curriculum: Controller-Shaped Grasping Behavior for Contact Force-Sensitive Manipulation

Ziyan Feng, Zizhao Yuan, Yulong Fu, Yuxin He, Zhiyuan Zhang, Zhengjie Zhang, Jinni Zhou, Renjing Xu, Qiang Nie

首次发表
浏览论文内容

中文总结 AI 辅助

针对力敏感操作中手动数据收集困难的问题,提出用25 Hz触觉反射控制器作为教师生成示范,训练无触觉策略,在ACT上实现95%稳定抓取,并揭示扰动抑制仍依赖控制器。

中文摘要 AI 辅助

机器人应如何学习操作那些亚牛顿接触力就可能造成不可逆损伤的易碎物体?现有的视觉-触觉策略学习通常将触觉感知视为额外的策略输入。然而,在直接接触力敏感操作中,瓶颈可能更早出现,即在数据收集阶段:手动夹爪控制过于延迟且粗糙,无法可靠地维持稳定抓取所需的窄力范围。因此,我们使用一个确定性的25 Hz触觉反射控制器作为收集时的教师,产生具有控制器塑造的抓取行为的示范,用于无触觉策略学习。在基于Transformer的动作分块(ACT)上,从反射塑造的示范训练出的策略恢复了教师的抓取轮廓,并在标称塑料杯任务上实现了95%的稳定抓取,显著优于视觉筛选的手动示范。同样的干预改善了π0.5在分布内的稳定性,并在未见过的纸杯变体上显示出有利的探索趋势。然而,在随机外部扰动下,反射数据的π0.5策略在仅策略试验中仍有45%失败,而部署时的反射仲裁器保留了所有抓取。这些结果揭示了触觉反馈在力敏感操作中的新角色:我们不是将触觉集成到策略中,而是将其用作收集时的教师,在示范中塑造抓取行为以供策略学习,而扰动抑制仍依赖于控制器,这揭示了无触觉策略的边界。

英文摘要

How should a robot learn to manipulate objects so fragile that sub-Newton contact forces can cause irreversible damage? Existing visuo-tactile policy learning typically treats tactile sensing as an additional policy input. In direct-contact force-sensitive manipulation, however, the bottleneck can arise earlier, during data collection: manual gripper control is too delayed and coarse-grained to reliably maintain the narrow force range required for stable grasping. We therefore use a deterministic 25 Hz tactile reflex controller as a collection-time teacher, producing demonstrations with controller-shaped grasping behavior for tactile-free policy learning. On Action Chunking with Transformers (ACT), policies trained from reflex-shaped demonstrations recover the teacher's grasping profile and achieve 95% stable grasps on the nominal plastic-cup task, substantially outperforming visually screened manual demonstrations. The same intervention improves in-distribution stability on $π_{0.5}$ and shows a favorable exploratory trend on an unseen paper-cup variant. Under randomized external disturbance, however, the reflex-data $π_{0.5}$ policy still fails in 45% of policy-only trials, whereas a deployment-time reflex arbiter retains all grasps. These results reveal a new role for tactile feedback in force-sensitive manipulation: rather than integrating tactile into the policy, we use it as a collection-time teacher that shapes grasping behavior in demonstrations for policy learning, while disturbance rejection remains controller-dependent, revealing the boundary of tactile-free policy.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑