arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11661cs.LGcs.AI

低交互秩学习:统一乘法双编码器头

Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads

Zijian Zhao, Sen Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出低交互秩学习框架,统一乘法双编码器头架构,解决其设计决策、可识别性等问题,经实验验证该框架可解释对比维度不可解释性并恢复真实模式。

中文摘要 AI 辅助

乘法双编码器网络通过将两个输入的独立编码的内积计算得到一对输入的实值输出。该架构已在算子学习、二分匹配、对比视觉-语言模型、检索及其他领域独立发展,但尚无统一理论指导基本设计决策:应表示多少种交互模式、如何归一化编码器,以及何时应避免使用该架构。我们通过引入低交互秩函数类提供了这样的基础,该类的固有复杂度由其交互谱衡量。在此框架内,近似误差分解为谱截断项和编码器实现项;样本复杂度由两个编码器复杂度之和而非乘积决定;基于谱衰减的可用性准则确定该架构何时能成功。同一框架揭示了一个核心可识别性问题:编码器仅由线性规范对称性定义,该对称性使学习到的坐标具有任意性。我们证明归一化是规范固定,而白化处理可将交互模式确定到排列和符号,从而解释了对比维度的不可解释性并提供了建设性解决方案。在合成核、算子学习和CLIP模型上的实验验证了理论预测:谱衰减率与预测缩放匹配,白化处理可恢复真实模式,独立训练的CLIP模型由单次旋转关联,经白化处理移除该旋转后可揭示可解释的概念轴。本文代码提供于此https://URL。

英文摘要

A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings. This architecture has been developed independently in operator learning, bipartite matching, contrastive vision-language models, retrieval, and other areas, yet no unified theory guides the basic design decisions: how many interaction modes to represent, how to normalize the encoders, and when the architecture should be avoided. We provide such a foundation by introducing the class of functions of low interaction rank, a class whose intrinsic complexity is measured by its interaction spectrum. Within this framework, approximation error decomposes into a spectral truncation term and an encoder-realization term; sample complexity is governed by the sum of the two encoder complexities rather than their product; and a usability criterion based on spectral decay determines when the architecture can succeed. The same framework exposes a central identifiability problem: the encoders are defined only up to a linear gauge symmetry that leaves the learned coordinates arbitrary. We show that normalization is gauge fixing and that whitening pins the interaction modes up to permutation and sign, thereby explaining the uninterpretability of contrastive dimensions and providing a constructive remedy. Experiments on synthetic kernels, operator learning, and CLIP models validate the theoretical predictions: spectral decay rates match the predicted scaling, whitening recovers the true modes, and independently trained CLIP models are related by a single rotation which, after removal by whitening, exposes interpretable concept axes. The code of this paper is provided at https://github.com/RS2002/Mul-Net .

发表机构

  • The Hong Kong University of Science and Technology(香港科技大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

↑