模态门控深度适配器:以精确保持的方式向冻结嵌入模型添加模态
Modality-Gated Deep Adapters: Adding a Modality to a Frozen Embedding Model with Exact Preservation
- Eximius Labs(Eximius实验室)
- Wabash College(瓦巴什学院)
- Skop Intelligence Co.(Skop智能公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出模态门控深度适配器,在不改变冻结嵌入模型现有输出的前提下添加新模态,通过按模态分组的瓶颈适配器实现零改变与零隔离,并在音频和热模态任务上显著提升检索性能。
AI中文摘要:
多模态嵌入模型被大规模部署:检索索引、基准测试结果和行为审计都依赖于基础模型的精确输出。使用现有的参数高效方法将此类模型扩展到新模态会静默改变这些输出;LoRA风格的适配无论权重是否合并都会重写文本路径,从而使所有存储的嵌入失效。我们提出模态门控深度适配器:瓶颈适配器附加到冻结的多模态嵌入LLM的每个解码器层,按模态分组打包,仅在编码自身模态时执行。结果是以对现有输出零改变的方式添加模态:没有包认领的输入逐位不变地通过基础模型自身的计算图,共同加载的包以精确零隔离矩阵组合。这两个性质均以命题形式陈述,在任意训练后成立而不仅仅在初始化时,推理时不需要任务标签或路由元数据,并通过发布检查点上的精确相等性测试验证。在一个冻结的2B基础模型上,音频包(以连接符令牌注入)将音频到文本R@10比同等训练的控制组提高+3.4至+5.4个百分点,每个种子均为正,并在十一倍数据下复现;热包复用基础模型自身冻结的视觉路径,在每个种子上约七倍地通过预注册的验收门槛,并将热到文本R@10从0.224提升至0.785。编码器交换定位缺失的能力:在CLAP风格比较中优于Whisper系列编码器的外部音频编码器在冻结LLM内部损失16个R@10点,因此能力属于层,正是门控适配器放置的位置。我们发布音频模型、热包以及训练、评估和不变性套件:模型见此URL,代码在GitHub上。
英文摘要:
Multimodal embedding models are deployed at scale: retrieval indices, benchmark results, and behavioral audits all depend on the base model's exact outputs. Extending such a model to a new modality with existing parameter-efficient methods silently changes those outputs; LoRA-style adaptation rewrites the text path whether or not the weights are merged, invalidating every stored embedding. We propose modality-gated deep adapters: bottleneck adapters attached to every decoder layer of a frozen multimodal embedding LLM, grouped into per-modality packs that execute only while their own modality is being encoded. The result is a modality added with zero change to existing outputs: inputs no pack claims traverse the base model's own computation graph, bit-for-bit unchanged, and co-loaded packs compose with an exact-zero isolation matrix. Both properties are stated as propositions, hold after arbitrary training rather than only at initialization, require no task labels or routing metadata at inference, and are verified by exact-equality tests on the released checkpoints. On one frozen 2B base, the audio pack (injected as connector tokens) improves audio-to-text R@10 by +3.4 to +5.4 points over an identically trained control, positive at every seed and reproduced at eleven times the data; the thermal pack, reusing the base's own frozen vision path, clears its pre-registered acceptance gate roughly sevenfold at every seed and lifts thermal-to-text R@10 from 0.224 to 0.785. An encoder swap locates the missing capacity: an external audio encoder that outranks Whisper-family encoders in CLAP-style comparisons loses by 16 R@10 points inside the frozen LLM, so the capacity belongs in the layers, exactly where the gated adapters place it. We release the audio model, the thermal pack, and the training, evaluation and invariance suites: models at huggingface.co/EximiusLabs, code on GitHub.