发表机构
H Company(H公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出NeoMME,一款单塔多模态原生多语言基础编码器,通过预训练后微调得到的检索模型在ViDoRe v3等基准测试中表现优异,还实现了高效的压缩与编码,相关资源已开源。
AI 中文摘要
多模态模型通常构建于为生成式视觉-语言建模设计的架构之上,一般将单独预训练的视觉编码器与因果语言模型相结合。视觉文档检索器(如ColPali)会将这些模型重新用作编码器,从而在非生成式任务中继承视觉语言模型(VLM)的参数与计算开销。本文提出NeoMME,这是一类包含2.6亿和8亿参数的多模态、多语言双向编码器,可在单个双向Transformer编码器中处理多语言文本与原始图像块。两个模型均从零开始预训练,采用掩码离散扩散文本目标函数,针对多模态示例以可见图像块作为条件。两者均支持16384个token的上下文,足以编码最多两张标准4K超高清图像。为验证其下游能力,我们采用联合训练的密集与后期交互头对NeoMME进行微调。在ViDoRe v3基准测试中,得到的NeoMME-Retriever 2.6亿参数模型以0.523的nDCG@10值,严格优于所有参数低于8亿的评估模型;NeoMME-Retriever 8亿参数模型则达到0.556的nDCG@10值。在NVIDIA L40S上采用匹配的2048×2048图像输入尺寸时,NeoMME-2.6亿参数模型对页面的编码吞吐量约为ColModernVBERT的2倍。分层token池化与非对称量化可将后期交互多模态文档嵌入压缩255倍,同时保留超过95%的基准nDCG@10值。我们将NeoMME贡献至Hugging Face Transformers,并在Apache 2.0许可下发布预训练主干与检索兼容检查点,链接为this https URL。
英文摘要
Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali repurpose these models as encoders, carrying over the parameter and compute overhead of a VLM for a non-generative task. We introduce NeoMME, a family of 260M and 800M-parameter Multimodal and Multilingual bidirectional Encoders that process multilingual text and raw image patches in a single bidirectional Transformer encoder. Both models are pretrained from scratch with a masked discrete-diffusion text objective, conditioned on visible image patches for multimodal examples. Both support a 16,384-token context, enough to encode up to two standard 4K UHD images. To demonstrate its downstream capabilities, we fine-tune NeoMME with jointly trained dense and late-interaction heads. On the ViDoRe v3 benchmark, the resulting NeoMME-Retriever 260M outperforms all evaluated models strictly below 800M parameters with 0.523 nDCG@10, while NeoMME-Retriever 800M reaches 0.556. At a matched 2048x2048 image input size on an NVIDIA L40S, NeoMME-260M encodes pages with about 2x the throughput of ColModernVBERT. Hierarchical token pooling and asymmetric quantization compress late-interaction multimodal document embeddings by 255x while preserving over 95% of baseline nDCG@10. We contribute NeoMME to Hugging Face Transformers and release the pretrained backbone and retrieval-compatible checkpoints under Apache 2.0 at https://hf.co/collections/Hcompany/neomme.