发表机构
Eximius Labs; Wabash College; Skop Intelligence Co.(卓越实验室; 瓦贝什学院; 斯科普智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在构建文本、图像、视频和音频统一嵌入空间。提出融合嵌入家族,第一代训练参数少的连接器,第二代添加模态门控深度适配器,无需成对视听数据实现音频-图像检索,探索设计空间并公开相关资源。
AI 中文摘要
一个涵盖文本、图像、视频和音频的单一嵌入空间能让一个索引服务用户的每一个查询。基于视觉语言主干构建的嵌入模型在文本/图像/视频检索基准测试中领先,但完全缺乏音频处理能力,而音频-文本检索则由无法处理其他模态的专业系统主导。我们提出了融合嵌入家族,它在参数从不更新的冻结视觉语言嵌入基础上添加音频:第一代(fusion-embedding-1)仅训练一个1640万个参数的连接器连接冻结的音频塔和冻结的基础,第二代(fusion-embedding-2)添加模态门控深度适配器(4420万个参数),其分支在文本、图像或视频输入上不执行,输出与发布的基础逐位相同,每次训练后验证。由于基础已经绑定了文本、图像和视频,仅将音频与文本对齐就能实现音频-图像检索,无需成对的视听训练数据。我们还通过可控的负面结果(用语言模型重写训练字幕、替换在排行榜上表现更强的音频塔、加宽连接器都会降低检索效果)和训练协议发现来探索其设计空间,这些发现有望应用于任何冻结的解码器-语言模型嵌入主干。两代模型在单个GPU上只需数小时就能训练完成。权重、代码和评估工具已公开发布。
英文摘要
A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead text/image/video retrieval benchmarks but lack audio entirely, while audio-text retrieval is led by specialist systems that serve no other modality. We present the Fusion Embedding family, which adds audio to a frozen vision-language embedding base whose parameters are never updated: generation 1 (fusion-embedding-1) trains only a 16.4M-parameter connector between a frozen audio tower and the frozen base, and generation 2 (fusion-embedding-2) adds modality-gated deep adapters (44.2M parameters) whose branch never executes on text, image, or video inputs: their outputs are bit-for-bit those of the released base, verified after every training run. Because the base already binds text, images, and video, aligning audio to text alone makes audio-image retrieval emerge, with zero paired audio-visual training data. Alongside the recipe we map its design space with controlled negative results (rewriting training captions with an LLM, substituting a leaderboard-stronger audio tower, and widening the connector each reduce retrieval) and with training-protocol findings that we expect to transfer to any frozen decoder-LM embedding backbone. Both generations train in hours on a single GPU. Weights, code, and the evaluation harness are openly released.
Comments23 pages, 5 figures. Models: https://huggingface.co/EximiusLabs/fusion-embedding-1-2b-preview and https://huggingface.co/EximiusLabs/fusion-embedding-2-2b-preview. Code: https://github.com/Eximius-Labs/fusion-embedding