重访分割:用于推理分割的工作记忆蒸馏
Revisit to Segment: Working Memory Distillation for Reasoning Segmentation
- Xiaohongshu(小红书)
- NTU(南洋理工大学)
- PKU(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出工作记忆蒸馏框架SWiM,通过利用模型自生成的工作记忆并蒸馏其指导,增强多模态大语言模型的推理分割能力,在基准上达到最优性能。
AI中文摘要:
多模态大语言模型(MLLMs)通过推理视觉内容并预测目标位置来进行图像分割。它们生成的响应包含推理轨迹和定位提议,当再次访问同一图像和查询时,这些可以作为工作记忆。我们的探索揭示,MLLMs 利用这种自生成的工作记忆作为上下文会受益,从而增强推理分割。受此发现启发,我们寻求通过蒸馏从重访先前尝试中获得的指导来增强骨干模型的推理分割能力,使其在推理时无论有无工作记忆都能受益。为此,我们提出了带有工作记忆的推理分割器(SWiM),一个用于推理分割的工作记忆蒸馏框架。具体来说,SWiM 基于分割质量选择 rollout 以构建工作记忆,并使用记忆条件模型作为教师。教师沿着学生生成的轨迹提供 token 级分布监督,而学生仅接收原始图像和查询。基于策略的自蒸馏和基于结果的强化学习的联合优化将工作记忆指导与对分割质量的直接反馈相结合。在推理分割基准上的大量实验表明,SWiM 实现了最先进的性能,验证了工作记忆蒸馏的有效性。
英文摘要:
Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Their generated responses contain reasoning traces and localization proposals that can serve as working memory when revisiting the same image and query. Our exploration reveals that MLLMs benefit from using this self-generated working memory as context, leading to enhanced reasoning segmentation. Motivated by this finding, we seek to strengthen the backbone model's reasoning segmentation capabilities by distilling the guidance gained from revisiting prior attempts, enabling it to benefit with or without working memory at inference time. To this end, we propose Reasoning Segmenter with Working Memory (SWiM), a working-memory distillation framework for reasoning segmentation. Specifically, SWiM selects rollouts based on segmentation quality to construct working memory and uses the memory-conditioned model as a teacher. The teacher provides token-level distributional supervision along student-generated trajectories, while the student receives only the original image and query. Joint optimization of on-policy self-distillation and outcome-based reinforcement learning combines working-memory guidance with direct feedback on segmentation quality. Extensive experiments on reasoning segmentation benchmarks demonstrate that SWiM achieves state-of-the-art performance, validating the effectiveness of working-memory distillation.