AI 中文总结
本文提出FIRM模型,通过预测视觉令牌内的r×r二值子单元掩码代码结合轻量连续渲染器优化边界,在5个遥感推理分割基准上取得领先结果,提升了细粒度分割性能。
AI 中文摘要
推理分割要求多模态大语言模型(MLLMs)将隐式指令转化为精确的像素级掩码。MLLMs将图像编码为视觉令牌,每个视觉令牌合并一组图像块。在遥感图像中,小目标、细结构及相邻实例可能占据同一视觉令牌的不同部分,为这类令牌分配单一二值掩码标签会丢失其内部空间结构,导致邻近目标合并、对象边界变粗糙。为弥合这一表示差距,本文提出FIRM(Fine-grained Intra-token Representation of Masks),即掩码的细粒度令牌内表示。对于每个视觉令牌,FIRM预测一个掩码代码,该代码指定r×r的二值子单元模式,而非单一前景/背景标签。给定MLLM识别的目标,可在一次掩码推理中预测完整的掩码代码网格;通过固定查找表将预测代码转换为离散子单元掩码,同时对代码分布求边际得到软结构场。为进一步恢复每个子单元内的细粒度边界,本文引入轻量连续渲染器,利用合并前的视觉特征和图像细节优化该结构场。在卫星和无人机图像的5个推理与指代表分割基准上,FIRM取得领先结果,包括在LaSeRS上达到70.5/80.5的gIoU/cIoU,在EarthReason上实现3.0点的平均提升。这些结果表明,显式表示令牌内掩码模式对细粒度MLLM分割具有重要价值。
英文摘要
Reasoning segmentation requires multimodal large language models (MLLMs) to translate implicit instructions into precise pixel-level masks. MLLMs encode an image as visual tokens, each of which merges a group of image patches. In remote sensing images, small targets, thin structures, and adjacent instances can occupy different parts of the same visual token. Assigning a single binary mask label to such a token loses its internal spatial structure, causing nearby targets to merge and object boundaries to become coarse. To bridge this representational gap, we introduce FIRM, a Fine-grained Intra-token Representation of Masks. For each visual token, FIRM predicts a mask code that specifies an $r\times r$ binary sub-cell pattern rather than a single foreground/background label. Given a target identified by the MLLM, the complete grid of mask codes is predicted in one mask pass. Fixed lookup converts the predicted codes into a discrete sub-cell mask, while marginalizing the code distribution yields a soft structural field. To further recover fine-grained boundaries within each sub-cell, we introduce a lightweight continuous renderer that refines this field using pre-merge visual features and image details. Across five reasoning and referring segmentation benchmarks on satellite and UAV images, FIRM achieves leading results, including $70.5/80.5$ gIoU/cIoU on LaSeRS and a $3.0$-point average gain on EarthReason. These results demonstrate the value of explicitly representing intra-token mask patterns for fine-grained MLLM segmentation.