Slot2Text:面向高效且空间可追溯的手术多模态大语言模型的以对象为中心的视觉分词
Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs
浏览论文内容
中文总结 AI 辅助
Slot2Text将视觉输入转为槽潜变量,推出双模式手术MLLM,在多基准上实现高效推理,Slot2Text-Fast大幅降本,Slot2Text-Reason支持可追溯空间推理。
中文摘要 AI 辅助
用于手术场景理解的多模态大语言模型(MLLM)通常会向语言模型中注入数百个密集视觉令牌,这会导致推理成本高昂,且生成答案的空间可追溯性受限。本文提出Slot2Text,这是一种双模式手术MLLM,它将视觉输入的密集表示替换为编码为槽潜变量的紧凑区域集合。与依赖视觉编码器与语言的对比对齐不同,Slot2Text将自监督视觉特征分组为少量区域——槽,这些槽被语言模型用作带区域标签的视觉令牌。Slot2Text-Fast利用槽前缀回答手术问题;Slot2Text-Reason还会识别并定位与推理相关的区域,将语言输出链接到对应的槽令牌、掩码或区域。在多个视觉问答和视觉定位基准上的实验表明,Slot2Text-Fast的性能与最先进的基线相当,但成本低得多,平均总令牌消耗减少91.8%,视觉前缀从1295个令牌降至47个(降幅达96.4%);Slot2Text-Reason则以额外的令牌和延迟为代价,换取明确的区域标识、位置和可追溯的空间证据。这些结果表明,紧凑的槽潜变量可作为手术MLLM的高效默认视觉接口,当需要更高空间可追溯性时,可调用基于基础的推理。
英文摘要
Multimodal large language models (MLLM) for surgical scene understanding typically inject hundreds of dense visual tokens into a language model, leading to costly inference and limited spatial traceability for generated answers. We present Slot2Text, a dual-mode surgical MLLM that replaces dense representations of visual input with a compact set of regions encoded as slot latents. Instead of relying on contrastive alignment of the visual encoder with language, Slot2Text groups self-supervised vision features into a few regions--slots that are consumed by the language model as area-labeled visual tokens. Slot2Text-Fast uses the slot prefix to answer surgical questions. Slot2Text-Reason also identifies and locates areas relevant for reasoning, linking language outputs to corresponding slot tokens, masks or regions. Experiments on multiple visual question answering and visual grounding benchmarks show that Slot2Text-Fast is competitive with state-of-the-art baseline at a much lower cost, reducing the average total token consumption by a 91.8\% and the visual prefix from 1,295 to 47 tokens (a 96.4\% reduction). Slot2Text-Reason trades additional tokens and latency for explicit area identities, locations, and traceable spatial evidence. These results establish compact slot latents as an efficient default visual interface for surgical MLLMs, with grounded reasoning invoked when greater spatial traceability is required.