arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SnapBench:面向移动交互的抓拍提问多模态检索基准测试

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

Zirong Chen, Fuda Ye, Kuan Zhang, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Jin Ma, Yongqi Zhang

arXiv 2608.29607首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Tencent Yuanbao; Tsinghua University; ARC Lab, Tencent; The University of Hong Kong(香港科技大学(广州); 腾讯元宝; 清华大学; 腾讯ARC实验室; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出首个鲁棒抓拍提问多模态检索配对基准SnapBench,评估16种多模态检索器,发现图像损坏严重影响检索,还提出自适应融合方法MOOR,凸显模态校准的必要性。

AI 中文摘要

移动人工智能作为视觉助手,支持用户拍摄物体照片并查询相关信息,抓拍提问检索是当前移动人工智能最常见的入口之一。然而抓拍的照片常存在模糊问题,而用户的文字提问可能简短或存在输入错误。现有基准仅针对干净输入进行测试,或未分离抓拍提问检索中的配对鲁棒性。因此,我们推出SnapBench,这是首个用于鲁棒抓拍提问多模态检索的配对基准,包含1145个查询、9085个图库项目,涵盖53种受控损坏条件并附带人工标注。我们评估了16种多模态检索器,包括双塔编码器和基于嵌入的视觉语言模型(VLMs)。结果显示,图像损坏会严重降低检索性能,而文本损坏主要影响纯文本检索,对联合检索的影响有限。干净的纯图像检索通常优于联合检索,这表明存在粗文本拖慢效应,以及在输入含噪时缺乏跨模态回退机制。SnapBench为评估抓拍提问场景中的鲁棒检索提供了受控测试平台。我们还提出了MOOR(Modality-anchored, Outlier-aware, Optimal Reweighting,即模态锚定、离群值感知的最优重加权),这是一种简单的自适应融合方法,凸显了抓拍提问检索中对可靠性感知模态校准的需求。

英文摘要

Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness in snap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-ask multimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, covering dual-tower encoders and embedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack of cross-modal fallback under noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further propose MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.

Comments37 pages. Yuanbao Technical Report. Accepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑