生成式嵌入基准:密集嵌入中保留了多少信息?
Generative Embedding Benchmark: How Much Information Survives in a Dense Embedding?
AI总结:
该研究提出Generative Embedding Benchmark(GEB),通过仅用嵌入和问题文本的解码器评估7种嵌入模型,发现生成式读出能揭示可分性评估未捕捉的信息瓶颈,视觉语言模型嵌入表现更优。
AI中文摘要:
嵌入已成为连接基础模型与下游系统的标准表示接口。大多数嵌入基准通过判别式任务或以嵌入空间可分性为中心的几何标准评估表示,但此类评估的优异性能无法证实压缩到嵌入中的内容是否仍可被下游生成器访问。为解决这一差距,我们提出生成式嵌入基准(Generative Embedding Benchmark,GEB):解码器仅使用冻结的嵌入和问题文本回答问题,无法访问原始图像或中间视觉特征,该读出方式下的答案质量可衡量生成式信息——即从嵌入中可恢复的与答案相关的内容。GEB包含精心整理的视觉问答数据集,其中包含1800项的开发集和900项的保留测试集,覆盖自然图像、场景文本和视觉文档。使用通用解码器和训练方案,我们在仅视觉模式和视觉-语言联合模式下评估了7种公开嵌入模型。在测试集上,仅视觉模式的得分范围为28.25至33.21;在图像-问题联合编码下,所有5种基于视觉语言模型(VLM)的嵌入模型得分更高,最佳得分达到65.56。匹配的嵌入也优于纯文本输入、零嵌入和打乱的嵌入。自然图像信息比场景文本或视觉文档信息更容易恢复,而可访问原始图像的Qwen3-VL-2B参考模型达到84.30。这些结果共同表明,生成式读出揭示了基于可分性的评估未捕捉到的信息瓶颈。
英文摘要:
Embeddings have emerged as a standard representational interface linking foundation models with downstream systems. Most embedding benchmarks assess representations through discriminative tasks or geometric criteria centered on separability in embedding space. However, strong performance on such evaluations does not establish whether content compressed into an embedding remains accessible to a downstream generator. To address this gap, we introduce the Generative Embedding Benchmark (GEB), in which a decoder answers questions using only a frozen embedding and question text, without access to the original image or intermediate visual features. Answer quality under this readout measures generative information: the answer-relevant content recoverable from an embedding. GEB includes a curated visual-question-answering dataset with a 1,800-item development split and a held-out 900-item test split covering natural images, scene text, and visual documents. Using a common decoder and training recipe, we evaluate seven public embedding models in visual-only and vision-language joint modes. On the test set, visual-only scores range from 28.25 to 33.21; with image-question joint encoding, all five VLM-based embedding models score higher, and the best reaches 65.56. Matched embeddings also outperform text-only inputs, zero embeddings, and shuffled embeddings. Natural-image information is much easier to recover than scene text or visual-document information, while a Qwen3-VL-2B reference with access to the original image reaches 84.30. Together, these results show that generative readout exposes information bottlenecks that separability-based evaluation does not capture.