arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24187cs.ROcs.CV

StenoVLA-3D:面向胃肠道狭窄导航的3D感知推理VLA

StenoVLA-3D: 3D-Aware Reasoning VLA for Navigation Through Gastrointestinal Stenoses

Tamima Tabassum, Yiming Huang, Tianchun Wu, Changjing Liu, Zhiqing Tang, Chikit Ng, Beilei Cui, Liangjing Shao, Jiewen Lai, Hongliang Ren

首次发表
浏览论文内容

中文总结 AI 辅助

针对内窥镜导航中纹理贫乏与病灶遗忘问题,提出3D感知VLA框架StenoVLA-3D,集成点图与时间状态分支,在测试中显著提升语义与动作准确率。

中文摘要 AI 辅助

自主内窥镜导航要求策略模型从纹理贫乏的单目观测中预测动作,做出安全的控制决策,并在病灶离开视野后保留其证据。现有的视觉-语言-动作(VLA)模型主要依赖视觉外观和短期上下文,限制了几何基础与片段级报告能力。我们提出StenoVLA-3D,一个用于狭窄区域导航的3D感知VLA框架。我们通过学习的几何门控融合将点图集成到Cosmos-Reason 2骨干网络中,并提出一个时间状态分支来建模遍历进度。我们的推理与动作骨干网络输出带有动作的接地推理,而专用头则估计狭窄形状并生成最终病灶报告。我们进一步引入EndoCausal,一个包含病灶标注、动作和时间接地推理的片段级数据集。在40个留出的记录测试片段上,StenoVLA-3D达到95.2%的语义准确率和83.4%的动作准确率。在物理3自由度内窥镜上,它在食管和结肠模型中(各36次试验)分别达到88.9%和77.8%的任务成功率,显著优于评估的基线方法。

英文摘要

Autonomous endoscopic navigation requires the policy model to predict actions from texture-poor monocular observations, make safe control decisions, and retain evidence of lesions after they leave the field of view. Existing vision-language-action (VLA) models primarily rely on visual appearance and short-term context, limiting geometric grounding and episode-level reporting. We introduce StenoVLA-3D, a 3D-aware VLA framework for navigating through stenotic regions. We integrate point-maps into the Cosmos-Reason 2 backbone through learned geometry-gated fusion, and also propose a temporal state branch to model traversal progress. Our reasoning-and-action backbone predicts grounded reasoning with actions, while dedicated heads estimate stenosis shape and generate the final lesion report. We further introduce EndoCausal, an episode-level dataset with lesion annotations, actions, and temporally grounded reasoning. On 40 held-out recorded test episodes, StenoVLA-3D reaches 95.2\% semantic accuracy and 83.4\% action accuracy. On the physical 3-DoF endoscope, it attains 88.9\% and 77.8\% task success in esophageal and colonic phantoms (36 trials each), substantially outperforming the evaluated baselines.

↑