PUMA:用于高效检索的通用多模态嵌入的事后稀疏化
PUMA: Post-Hoc Sparsification of Universal Multimodal Embeddings for Efficient Retrieval
AI总结:
针对通用多模态嵌入的高存储与推理成本问题,提出无需微调主干网络的PUMA稀疏自编码器方案,在五个基准上验证其可实现8-16倍存储压缩与最高25倍检索加速,且多数数据集性能与密集检索相当或更优。
AI中文摘要:
通用多模态嵌入器支持跨文本、图像及组合查询的检索,但其密集表示会产生高内存与推理成本。事后稀疏化可降低这些成本,但在多模态检索领域仍未得到充分探索。我们提出PUMA,一种稀疏自编码器方案,可在不微调主干网络的前提下将通用多模态嵌入映射为紧凑稀疏码:预训练阶段保留密集点积几何结构,之后对稀疏编码器进行微调以适配检索任务。我们在涵盖文本到图像及组合图像检索的五个基准上进行评估,在Qwen3-VL-Embedding-2B模型上,PUMA在五个数据集中的四个上与密集检索的性能无统计差异或实现了性能提升。我们进一步识别出事后稀疏化的两种失效模式:TopK前支持不足及检索对齐的主动支持缺失。PUMA将向量存储减少8-16倍(FP32),在更大候选池上比精确密集评分快达25倍,可实现高效多模态检索。
英文摘要:
Universal multimodal embedders enable retrieval across text, image, and combined queries, but their dense representations incur high memory and inference costs. Post-hoc sparsification could reduce these costs but remains underexplored for multimodal retrieval. We introduce PUMA, a sparse autoencoder recipe that maps universal multimodal embeddings to compact sparse codes without retraining the backbone: a pretraining stage preserves dense dot-product geometry, after which the sparse encoder is fine-tuned for retrieval. We evaluate on five benchmarks covering text-to-image and composed image retrieval. On Qwen3-VL-Embedding-2B, PUMA is statistically indistinguishable from or improves over dense retrieval on four of five datasets. We further identify two failure modes of post-hoc sparsification: insufficient pre-TopK support and retrieval-misaligned active support. PUMA reduces vector storage by 8-16x (FP32) and is up to 25x faster than exact dense scoring on larger candidate pools, enabling efficient multimodal retrieval.