DeCo:面向多任务视觉定位的高效解耦-耦合学习
DeCo: Efficient Decouple-to-Couple Learning for Multi-Task Visual Grounding
浏览论文内容
中文总结 AI 辅助
DeCo提出解耦-耦合两阶段学习框架,通过任务感知语义解耦与混合先验耦合,在冻结多模态编码器上实现多任务视觉定位的最先进性能。
中文摘要 AI 辅助
多任务视觉定位要求模型联合理解语言语义,并执行准确的视觉定位和分割。尽管多模态大语言模型取得了成功,但有效地将其适配到多个定位目标仍然具有挑战性。现有方法通常通过共享表示来强制任务协作,却忽视了任务导向特征兴趣之间的内在冲突。本文提出了DeCo,一个高效的“解耦到耦合”学习框架,通过两阶段范式解决这一困境:先进行任务特定表示解耦,再进行互补先验耦合。具体而言,我们首先提出任务感知语义解耦(TSD),在显著词级引导下将共享视觉线索路由到各个特征中,减轻定位与分割之间的表示干扰。此外,我们观察到分割由于密集监督自然提供了信息丰富的定位先验。基于此洞察,我们引入混合先验耦合(HPC),将句子级语义先验与掩码导出的空间先验相结合以增强定位。基于冻结的多模态编码器,DeCo仅需轻量级可训练参数,同时实现跨多个定位目标的强泛化。在RefCOCO/+、G-Ref、ReferIt、Flickr、DIOR-RSVG、SARVG1.0、RRSIS-D、RIS-LAD和RefDIOR上的大量实验表明,DeCo在自然和遥感基准上均达到最先进性能。代码和模型可在该https URL获取。
英文摘要
Multi-task visual grounding requires models to jointly understand linguistic semantics and perform accurate visual localization and segmentation. Despite the success of multimodal large language models, effectively adapting them to multiple grounding objectives remains challenging. Existing methods commonly enforce task cooperation through shared representations, while overlooking the intrinsic conflict between task-oriented feature interests. In this paper, we introduce $\textbf{DeCo}$, an efficient $\textbf{De}$couple-to-$\textbf{Co}$uple learning framework that resolves this dilemma through a two-stage paradigm: task-specific representation decoupling followed by complementary prior coupling. Specifically, we first propose Task-aware Semantic Decoupling (TSD) to route shared visual cues into individual features under salient word-level guidance, alleviating representation interference between localization and segmentation. Furthermore, we observe that segmentation naturally provides informative localization priors due to dense supervision. Based on this insight, we introduce Hybrid Prior Coupling (HPC), which integrates sentence-level semantic prior with mask-derived spatial prior for enhanced grounding. Built upon a frozen multimodal encoder, DeCo requires lightweight trainable parameters while achieving strong generalization across multiple grounding objectives. Extensive experiments on RefCOCO/+, G-Ref, ReferIt, Flickr, DIOR-RSVG, SARVG1.0, RRSIS-D, RIS-LAD, and RefDIOR demonstrate that DeCo achieves state-of-the-art performance on both natural and remote sensing benchmarks. The code and models are available at https://github.com/xiaoqiang-lu/DeCo.
发表机构
- Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education(教育部智能感知与图像理解重点实验室)
- School of Artificial Intelligence, Xidian University(西安电子科技大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。