EndoLIFT:用于双向内镜控制的语言消歧潜在条件修正流
EndoLIFT: Language-Disambiguated Latent-Conditioned Rectified Flow for Bidirectional Endoscopic Control
浏览论文内容
中文总结 AI 辅助
EndoLIFT是一种视觉-语言-动作策略,通过结合语言意图条件与潜在条件修正流,解决双向内镜控制的意图混叠问题,在导航准确率、错误方向减少及多 phantom 试验中均取得显著提升。
中文摘要 AI 辅助
常规胃肠道内镜检查本质上是双向的:器械被推进以到达目标解剖部位,随后被撤回或反转以进行检查,而外部提示可能要求更早反转。当请求的阶段在视觉场景变化前改变时,几乎相同的观察结果可能需要相反的轴向动作。我们将双向内镜控制中的这种歧义识别并形式化为意图混叠。我们提出EndoLIFT(带轨迹潜在变量的内镜语言指令流),这是一种视觉-语言-动作策略,结合了显式基于语言的意图条件和潜在条件修正流动作专家。该策略接收RGB图像、语言指令和前一动作状态;一个32维变分轨迹潜在变量随机条件化连续动作块的生成。受控的相同观察指令交换表明,语言选择轴向模式,与是否存在轨迹潜在变量无关。相对于没有潜在条件的匹配模型,EndoLIFT将导航方向准确率提高了11.1个百分点,并减少了83%的错误方向推进。架构控制的1位模式标志参考表现出较弱的规范锚切换,而EndoLIFT在44个未见过的语言变体上保持了82.8%的意图跟随准确率。在闭环评估中,EndoLIFT在已见结肠 phantom 和未见的肺、胃 phantom 上,比无VTL的EndoLIFT整体成功率提高了30个百分点,并完成了10/10的离体猪气管试验。这些结果将基于语言的意图选择与轨迹潜在变量对方向正确性和稳健撤回的贡献分离开来。
英文摘要
Routine gastrointestinal endoscopy is intrinsically bidirectional: the instrument is advanced to reach target anatomy and later withdrawn or retroflexed for inspection, while an external cue may require earlier reversal. When the requested phase changes before the visual scene does, nearly identical observations can require opposite axial actions. We identify and formalize this ambiguity in bidirectional endoscopic control as intent aliasing. We propose EndoLIFT (Endoscopic Language-Instruction Flow with Trajectory Latents), a vision-language-action policy that combines explicit language-based intent conditioning with a latent-conditioned rectified-flow action expert. The policy receives RGB, a language instruction, and the previous-action state; a 32-D variational trajectory latent stochastically conditions continuous action-chunk generation. Controlled same-observation instruction swaps establish that language selects the axial mode, independently of whether the trajectory latent is present. Relative to the matched model without latent conditioning, EndoLIFT improves navigation-direction accuracy by 11.1 percentage points and reduces wrong-direction advance by 83\%. An architecture-controlled 1-bit mode-flag reference exhibits weaker canonical-anchor switching, while EndoLIFT retains 82.8\% intent-following accuracy across 44 held-out linguistic variants. In closed-loop evaluation, EndoLIFT improves overall success by 30 percentage points over EndoLIFT w/o VTL on both the seen colon phantom and the unseen lung and stomach phantoms, and completes 10/10 ex-vivo porcine-trachea trials. These results separate language-based intent selection from the trajectory latent's contribution to directional correctness and robust retraction.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学附属第六医院)
机构由 AI 辅助整理,请以论文原文为准。