PixVL:通过统一掩码-文本一致性循环进行像素级多模态大语言模型的自监督训练
PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask--Text Consistency Cycle
AI总结:
PixVL作为自监督后训练框架,通过统一掩码-文本一致性循环及语义验证等策略,解决像素级MLLMs的优化干扰问题,同时提升区域理解与分割任务性能。
AI中文摘要:
近期研究开发了支持区域分割和区域理解的像素级多模态大语言模型(MLLMs),将多模态交互从整幅图像扩展到特定对象和区域。然而这些方法面临两个核心挑战:一是高质量掩码-文本对稀缺,大量掩码标注缺乏对应的语言监督;二是监督格式和学习信号密度的差异,导致区域分割与区域理解之间出现优化干扰。为解决这些挑战,我们提出PixVL,一种自监督后训练框架,引入统一的掩码-文本一致性循环,使像素级MLLMs能生成并自我验证区域描述,从未标注数据中学习。我们发现仅基于几何重建的直接循环不可靠,因为重分割交并比(IoU)无法忠实反映语义质量和指代充分性。因此PixVL引入感知混淆项的语义验证,利用模型在高度相似候选区域中正确选择目标时的置信度,对错误选择分配零奖励;同时,PixVL使用时间分离的视频帧或几何变换的图像视图执行跨视图验证,防止循环学习坍缩为位置和形状捷径。最后,质量耦合双向学习策略用最高奖励的描述引导文本到掩码学习,该策略将区域理解和区域分割从竞争任务转变为相互生成器和验证器。实验表明PixVL同时提升了区域理解任务和分割任务的性能。
英文摘要:
Recent studies develop pixel-level multimodal large language models (MLLMs) that support both Region Segmentation and Region Understanding, extending multimodal interaction from whole images to specific objects and regions. However, these methods face two fundamental challenges. First, the scarcity of high-quality mask--text pairs leaves abundant mask annotations without corresponding language supervision. Second, discrepancies in supervision formats and learning-signal densities induce optimization interference between Region Segmentation and Region Understanding. To address these challenges, we propose PixVL, a self-supervised post-training framework that introduces a unified Mask--Text Consistency Cycle, enabling pixel-level MLLMs to generate and self-verify regional descriptions and learn from unlabeled data. We found that direct cycle based solely on geometric reconstruction is unreliable because re-segmentation IoU does not faithfully reflect the semantic quality and referring sufficiency. PixVL therefore introduces confuser-aware semantic verification, which uses the model's confidence when it correctly chooses the target among highly similar candidate regions, and assigns zero reward to an incorrect choice. Meanwhile, PixVL performs cross-view verification using temporally separated video frames or geometrically transformed image views, preventing cyclic learning from collapsing to positional and shape shortcuts. Finally, a quality-coupled bidirectional learning strategy uses the highest-reward description to guide Text-to-Mask learning. This strategy transforms Region Understanding and Region Segmentation from competing tasks into mutual generators and verifiers. Experiments demonstrate that PixVL improves both region understanding task and segmentation task.