基于跨模态注意力先验的文本引导视觉依赖图学习
Text-Guided Visual Dependency Graph Learning with Cross-Modal Attention Priors
浏览论文内容
中文总结 AI 辅助
提出CM-GLasso框架,结合文本可视化、交叉注意力蒸馏和联合ADMM,在八个基准上取得最优分类与分割性能,可生成可解释的稀疏条件依赖图。
中文摘要 AI 辅助
从多模态视觉-语言特征中估计可解释的条件依赖结构,仍是一个尚未充分探索的问题。我们提出CM-GLasso(跨模态图套索,Cross-Modal Graphical Lasso),这一框架将视觉-语言表征学习与稀疏高斯图模型相连接。CM-GLasso包含三个关键组件:(i) 文本可视化策略,将类别属性描述渲染为图像,并通过与处理自然图像相同的SigLIP-2视觉编码器进行处理,在共享特征坐标系中产生以原型索引的补丁级注意力足迹;(ii) 交叉注意力蒸馏机制,将高维补丁浓缩为一小部分语义图节点,其注意力足迹相似度为非均匀L1惩罚提供跨模态结构先验;(iii) 联合ADMM公式,在单个凸目标内估计共享和类别特定的精度分量,避免先估计再分解单独的类别级图的需求。所学的稀疏图拓扑直接支持无参数、基于精度的分类规则,以及轻量级的拓扑感知分割头。在八个基准上的大量实验表明,与强大的基于特征的和任务特定的基线相比,CM-GLasso取得了有竞争力或更优的性能。在匹配的受控协议下,它在VOC(74.75%)和ADE20K(64.01%)的受控基线中达到了最高的平均分类准确率(91.97%)和最高的分割mIoU,同时还产生了具有类别特定分解的显式稀疏条件依赖图。
英文摘要
Estimating interpretable conditional-dependence structures from multimodal visual-linguistic features remains largely unexplored. We propose CM-GLasso (Cross-Modal Graphical Lasso), a framework that bridges vision-language representation learning and sparse Gaussian Graphical Models. CM-GLasso introduces three key components: (i) a text visualization strategy that renders class-attribute descriptions as images and processes them through the same SigLIP-2 vision encoder as natural images, yielding prototype-indexed patch-level attention footprints in a shared feature coordinate system; (ii) a cross-attention distillation mechanism that condenses high-dimensional patches into a small set of semantic graph nodes, whose attention-footprint similarities yield cross-modal structural priors for non-uniform L1 penalization; (iii) a joint ADMM formulation that estimates shared and class-specific precision components within a single convex objective, avoiding the need to first estimate and then decompose separate class-wise graphs. The learned sparse graph topologies directly support a parameter-free, precision-based classification rule and a lightweight topology-aware segmentation head. Extensive experiments on eight benchmarks demonstrate that CM-GLasso achieves competitive or superior performance compared with strong feature-based and task-specific baselines. Under the matched controlled protocol, it attains the highest average classification accuracy (91.97%) and the highest segmentation mIoU among the controlled baselines on VOC (74.75%) and ADE20K (64.01%), while also yielding explicit sparse conditional-dependence graphs with common-specific decomposition.