arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21365cs.RO

MicroHookACT:单目显微视觉引导的柔性微电极钩取柔顺视触觉策略

MicroHookACT: Monocular Microscopic Vision Guided Visuomotor Policy for Flexible Microelectrode Hooking

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • State Key Laboratory of Brain Cognition and Brain-Inspired Intelligence Technology(脑认知与脑启发智能技术国家重点实验室)
  • Institute of Semiconductors, Chinese Academy of Sciences(中国科学院半导体研究所)

机构由 AI 辅助整理,请以论文原文为准。

Yitong Chen, Fangbo Qin, Yang Wang, Ruihua Hu, Kui Zhang, Shan Yu

AI总结:

本文提出MicroHookACT,一种基于模仿学习的视触觉策略,利用单目显微视觉和散焦线索实现柔性微电极自动钩取,在五种难度设置下达到96.7%的成功率。

AI中文摘要:

自动针环钩取是柔性微电极(FME)植入中的关键步骤。本文提出MicroHookACT,一种基于模仿学习的视触觉策略,用于在单目显微视觉下实现自动三维钩取。首先,一种单向钩取策略利用散焦线索和光轴引导,实现无需触觉探查的精确对准和富含接触的穿线操作。其次,一个基于冻结ViT骨干网络的动作监督对象注意力模块,直接从人类演示中学习聚焦于微针尖端和微环,无需手动视觉标注进行训练。第三,根据预测的动作进度,动态加权以注意力为中心的全局粗特征和局部细特征,使单一ACT策略能够适应整个操作过程中变化的散焦模糊和视觉需求。在实验中,视触觉策略在60个人类演示上训练,并在五种不同难度的设置下进行评估。我们的MicroHookACT框架实现了最高的总体成功率96.7%,平均执行时间为11.5秒。这些结果展示了视触觉策略学习在不同操作条件下实现微米级控制的潜力。

英文摘要:

Automated needle-loop hooking is a critical step in flexible microelectrode (FME) implantation. This paper presents MicroHookACT, an imitation learning-based visuomotor policy for automated 3D hooking under monocular microscopic vision. First, a unidirectional hooking strategy exploits defocus cues and optical-axis guidance to enable palpation-free precise alignment and contact-rich threading. Second, an action-supervised object attention module built on a frozen ViT backbone learns to focus on the micro-needle tip and micro-loop directly from human demonstrations, without requiring manual visual annotations for training. Third, attention-centered global coarse and local fine features are dynamically weighted according to predicted action progress, enabling a single ACT policy to adapt to changing defocus blur and visual requirements throughout the operation. In the experiments, visuomotor policies were trained on 60 human demonstrations and evaluated under five setups with varying difficulties. Our MicroHookACT framework achieved the highest overall success rate of 96.7\% with an average execution time of 11.5 s. These results demonstrate the potential of visuomotor policy learning for micron-level control under varying operating conditions.

↑