StenoVLA-3D:面向胃肠道狭窄导航的3D感知推理VLA
StenoVLA-3D: 3D-Aware Reasoning VLA for Navigation Through Gastrointestinal Stenoses
浏览论文内容
中文总结 AI 辅助
针对内窥镜导航中纹理贫乏与病灶遗忘问题,提出3D感知VLA框架StenoVLA-3D,集成点图与时间状态分支,在测试中显著提升语义与动作准确率。
中文摘要 AI 辅助
自主内窥镜导航要求策略模型从纹理贫乏的单目观测中预测动作,做出安全的控制决策,并在病灶离开视野后保留其证据。现有的视觉-语言-动作(VLA)模型主要依赖视觉外观和短期上下文,限制了几何基础与片段级报告能力。我们提出StenoVLA-3D,一个用于狭窄区域导航的3D感知VLA框架。我们通过学习的几何门控融合将点图集成到Cosmos-Reason 2骨干网络中,并提出一个时间状态分支来建模遍历进度。我们的推理与动作骨干网络输出带有动作的接地推理,而专用头则估计狭窄形状并生成最终病灶报告。我们进一步引入EndoCausal,一个包含病灶标注、动作和时间接地推理的片段级数据集。在40个留出的记录测试片段上,StenoVLA-3D达到95.2%的语义准确率和83.4%的动作准确率。在物理3自由度内窥镜上,它在食管和结肠模型中(各36次试验)分别达到88.9%和77.8%的任务成功率,显著优于评估的基线方法。
英文摘要
Autonomous endoscopic navigation requires the policy model to predict actions from texture-poor monocular observations, make safe control decisions, and retain evidence of lesions after they leave the field of view. Existing vision-language-action (VLA) models primarily rely on visual appearance and short-term context, limiting geometric grounding and episode-level reporting. We introduce StenoVLA-3D, a 3D-aware VLA framework for navigating through stenotic regions. We integrate point-maps into the Cosmos-Reason 2 backbone through learned geometry-gated fusion, and also propose a temporal state branch to model traversal progress. Our reasoning-and-action backbone predicts grounded reasoning with actions, while dedicated heads estimate stenosis shape and generate the final lesion report. We further introduce EndoCausal, an episode-level dataset with lesion annotations, actions, and temporally grounded reasoning. On 40 held-out recorded test episodes, StenoVLA-3D reaches 95.2\% semantic accuracy and 83.4\% action accuracy. On the physical 3-DoF endoscope, it attains 88.9\% and 77.8\% task success in esophageal and colonic phantoms (36 trials each), substantially outperforming the evaluated baselines.