发表机构
North China University of Technology; Central South University; The University of Sydney; Sun Yat-sen University; National University of Singapore(北方工业大学; 中南大学; 悉尼大学; 中山大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Hub-谱激活(HSA)方法,通过闭式谱分析激活冻结表示中的潜在多模态知识,无需监督或训练,显著提升跨模态检索与分类性能。
AI 中文摘要
多模态表示学习旨在为跨模态检索和知识迁移寻找共享表示。基于Hub的绑定降低了成对监督的成本,但独立的Hub连接无法在没有直接联合训练的情况下保证模态间的可靠对齐。我们提出了Hub-谱激活(HSA),一种闭式方法,用于在冻结表示中恢复并激活潜在多模态知识的Hub可读组件。我们将该知识形式化为源诱导的跨模态依赖,并刻画由两个训练过的Hub边的二阶统计量决定的组件。在二阶源模型下,我们建立了完整恢复源诱导关系的条件,并将Hub可读组件的维度上界界定为Hub协方差秩。HSA组合并标准化Hub边统计量,提取成对谱方向,并将可靠性加权匹配证据与源门控候选解析相结合,用于双向检索和原型分类。HSA不需要目标对监督、梯度优化或骨干网络更新。在ImageBind和LanguageBind上的19个检索关系和11个原型分类关系中,HSA将平均双向Recall@10从18.27%提升至31.15%,平均宏Top-1准确率从29.01%提升至52.43%。受控分析进一步识别出有效的Hub边对应和主导谱方向是检索增益的关键来源,证明了潜在多模态知识超越原生相似度分数的实用性。代码和模型可在该https URL公开获取。
英文摘要
Multimodal representation learning seeks shared representations for cross-modal retrieval and knowledge transfer. Hub-based binding reduces pairwise supervision costs, but separate hub connections cannot guarantee reliable alignment between modalities without direct joint training. We introduce Hub-Spectral Activation (HSA), a closed-form method for recovering and activating the hub-readable component of latent multimodal knowledge in frozen representations. We formalize this knowledge as source-induced cross-modal dependence and characterize the component determined by the second-order statistics of two trained hub edges. Under a second-order source model, we establish conditions for exact recovery of the complete source-induced relation and bound the dimension of its hub-readable component by the hub covariance rank. HSA composes and standardizes hub-edge statistics, extracts paired spectral directions, and combines reliability-weighted matching evidence with source-gated candidate resolution for bidirectional retrieval and prototype classification. HSA requires no target-pair supervision, gradient optimization, or backbone updates. Across 19 retrieval and 11 prototype-classification relations on ImageBind and LanguageBind, HSA raises mean bidirectional Recall@10 from 18.27% to 31.15% and mean macro Top-1 accuracy from 29.01% to 52.43%, respectively. Controlled analyses further identify valid hub-edge correspondence and leading spectral directions as key sources of retrieval gains, demonstrating the utility of latent multimodal knowledge beyond native similarity scores. Code and models are publicly available at https://github.com/Luo1Yan/HSA.
Comments30 pages, 9 figures, including appendices