发表机构
The Hong Kong University of Science and Technology (HKUST)(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
STAMPlus通过结构化全掩码预测解耦自回归对话与非自回归掩码预测,解决了MLLM分割的三难问题,在提升性能的同时降低延迟,实现多类开放词汇等分割任务的SOTA表现。
AI 中文摘要
基于多模态大型语言模型(MLLM)的分割面临核心的分割三难问题:高分割性能、保留对话能力以及快速推理。嵌入预测方法可能会通过像素级目标破坏语言建模,而下一个令牌生成对于密集掩码而言效率低下。我们提出全掩码预测,将自回归对话与非自回归掩码预测解耦。其二元实例化版本STAMP(Simultaneous Textual All-Mask Prediction,同步文本全掩码预测)会生成词汇内的<SEG>触发符,融合与图像对齐的掩码令牌及对应图像块特征,并使用混合注意力在一次前向传播中将所有令牌分类为前景或背景。这使其结合了强 referring 分割与推理分割能力,同时保留多模态能力且推理高效。然而,二元掩码无法在不重复特定目标预测的情况下保留多个语义或实例身份,因此我们提出结构化全掩码预测并开发STAMPlus。它生成带有显式ID和可选边界框的目标列表,将这些ID绑定到共享的多类掩码空间,并在一次非自回归前向传播中联合预测所有目标。单个统一检查点保留了STAMP的referring与推理能力,同时扩展到开放词汇语义分割、实例感知分割以及遥感小目标分割,其中高分辨率掩码令牌缩放保留了更精细的空间证据。在这些设置中,STAMPlus实现了SOTA分割性能,保留了通用多模态指令跟随能力,并将12类延迟从重复STAMP推理的13.50秒降至5.16秒。进一步分析显示,准确的目标线索可提升分割效果,且学习到的空间接地有益于二次查看推理。总体而言,STAMPlus在超越单目标预测的情况下解决了该三难问题。
英文摘要
MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt language modeling through pixel-level objectives, whereas next-token generation is inefficient for dense masks. We propose All-Mask Prediction, decoupling autoregressive dialogue from non-autoregressive mask prediction. Its binary instantiation, STAMP (Simultaneous Textual All-Mask Prediction), emits an in-vocabulary <SEG> trigger, fuses image-aligned mask tokens with corresponding patch features, and uses hybrid attention to classify all tokens as foreground or background in one pass. It thereby combines strong referring and reasoning segmentation with preserved multimodal ability and efficient inference. However, binary masks cannot retain multiple semantic or instance identities without repeated target-specific predictions. We therefore propose Structured All-Mask Prediction and develop STAMPlus. It generates a target list with explicit IDs and optional boxes, binds these IDs to a shared multi-class mask space, and jointly predicts all targets in one non-autoregressive pass. A single unified checkpoint retains STAMP's referring and reasoning capabilities while extending to open-vocabulary semantic, instance-aware, and remote-sensing small-target segmentation, where high-resolution mask-token scaling preserves finer spatial evidence. Across these settings, STAMPlus achieves state-of-the-art segmentation performance, preserves general multimodal instruction following, and reduces 12-category latency from 13.50s for repeated STAMP inference to 5.16s. Further analyses show that accurate target cues improve segmentation and learned spatial grounding benefits look-twice reasoning. Overall, STAMPlus resolves the trilemma beyond single-target prediction.