arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ColoACT:用于自推进内窥镜机器人平滑自主结肠导航的多线索动作分块

ColoACT: Multi-Cue Action Chunking for Smooth Autonomous Colon Navigation on a Self-Propelled Endoscopic Robot

Jian Hu, Shujing He, Leixin Chang, Zongze Li, Ding Huang, Chaoyang Shi, Chengzhi Hu

arXiv 2610.01258首次发表:更新:

发表机构

Southern University of Science and Technology; Zhejiang University; Tianjin University(南方科技大学; 浙江大学; 天津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自推进内窥镜机器人,提出ColoACT系统,采用RGB-D-E动作分块Transformer策略,通过多线索增强和时间集成实现平滑自主结肠导航,在离体猪结肠中达到85.4%和72.5%的直段和弯曲段成功率。

AI 中文摘要

自主结肠镜导航可以减少操作者负担以及环形成或组织创伤的风险,但由于可变形解剖结构、弱纹理和高光内镜视觉以及富含接触的粘弹性相互作用,仍然具有挑战性。现有方法要么依赖几何驱动的流程,这些流程高效且可解释,但由于手工设计的特征和切换逻辑而脆弱;要么采用基于学习的策略,其推断的深度/几何在弱纹理和高光下可能变得时间上不一致或过度平滑,而模拟训练的变体(例如深度强化学习)可能进一步遭受模拟到现实的差距。我们提出了ColoACT,一种自主导航系统,集成了基于RGB-D-E的动作分块Transformer策略(ColoACT策略),用于紧凑型自推进斜齿轮内窥镜机器人(BGER)。ColoACT策略通过估计的相对深度和基于梯度的伪高程图增强RGB,以增强褶皱脊显著性和其他高频几何线索,并通过预测重叠动作块并经由时间集成融合,实现对BGER的平滑连续控制。在不同的离体猪结肠(约60厘米)中,我们的系统在直段和弯曲段分别实现了85.4%和72.5%的成功率,在90度转弯中实现70%的成功率,在双弯序列中实现60%的成功率,并在具有挑战性的三弯段中进一步证明了可行性。项目页面可在以下网址获取:this https URL。

英文摘要

Autonomous colonoscopic navigation can reduce operator burden and the risk of loop formation or tissue trauma, but remains challenging due to deformable anatomy, weak-texture and specular endoscopic visuals, and contact-rich viscoelastic interactions. Existing methods either rely on geometry-driven pipelines, which are efficient and interpretable yet brittle due to manually engineered features and switching logic, or adopt learning-based policies, whose inferred depth/geometry can become temporally inconsistent or overly smooth under weak texture and specular highlights while simulation-trained variants (e.g., deep reinforcement learning) may further suffer from a sim-to-real gap. We propose ColoACT, an autonomous navigation system that integrates an RGB-D-E based Action Chunking Transformer policy (ColoACT policy) for a compact self-propelled Bevel-Gear-Based Endoscopic Robot (BGER). The ColoACT policy augments RGB with estimated relative depth and a gradient-based pseudo-elevation map to enhance fold-ridge saliency and other high-frequency geometric cues, and enables smooth continuous control of the BGER by predicting overlapping action chunks and fusing them via temporal ensembling. In different \textit{ex-vivo} porcine colons (approximately 60 cm), our system achieves success rates of 85.4\% and 72.5\% in straight and curved segments, respectively, and achieves 70\% success in 90-degree turns and 60\% in double-bend sequences, with feasibility further demonstrated in challenging triple-bend segments. The project page is available at: https://Adamhu1.github.io/ColoACT/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑