arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

前沿大语言模型能否媲美原生多模态嵌入?在难负例文本到图像检索上的对比

Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval

Archan Dutta, Vyanktesh Kanungo

arXiv 2608.11343首次发表:更新:

发表机构

Westcliff University(韦克利夫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在Flickr30k数据集上对比Gemini Embedding 2等原生多模态嵌入与GPT-4.1、Claude Sonnet 4.6等前沿LLM的难负例文本到图像检索性能,发现二者表现相当,且多模态嵌入更适配低延迟应用

AI 中文摘要

跨文本、图像、视频、音频等不同媒体类型的多模态检索与分类,传统上依赖通过对比学习对齐视觉与文本表示的双编码器模型。2026年3月谷歌发布的Gemini Embedding 2是其首款原生多模态嵌入模型,可将文本、图像、视频、音频及文档映射到单一共享空间,加剧了多模态检索系统的竞争。与此同时,前沿大语言模型(LLMs)已展现出强大的视觉理解能力,引发了“它们能否作为有效零样本排序器”的疑问。本研究首次在Flickr30k数据集上直接对比原生多模态嵌入与基于LLM的视觉排序性能,观察到GPT-4.1与Claude Sonnet 4.6的表现与Gemini Embedding 2相当;此外,嵌入预计算完成后,多模态嵌入更适合低延迟应用场景。

英文摘要

Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 release of Gemini Embedding 2, Google's first natively multimodal embedding model to map text, images, video, audio, and documents into a single shared space, raises competition among multimodal retrieval systems. Simultaneously, frontier Large language models (LLMs) have also demonstrated strong visual understanding, raising the question of whether they can serve as effective zero-shot rankers. Our study provides the first direct comparison of native multimodal embeddings against LLM-based visual ranking on Flickr30k. We observe that GPT-4.1 and Claude Sonnet 4.6 perform on par with Gemini Embedding 2. Additionally, once embeddings are precomputed, multimodal embeddings are better suited for low-latency applications.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑