arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15695cs.AI

NoteVQA:基于人类社区真实问题的视觉语言模型基准测试

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

  • AllSpark Team(AllSpark团队)

机构由 AI 辅助整理,请以论文原文为准。

Haonan Jiang, Guojian Zhan, Jiancong Xie, Shijun Wan, Dongiia Zhao, Cheng Chen, Yahui Liu, Chuan Mu

AI总结:

针对日常视觉问题评估VLM的空白,构建NoteVQA基准(252条目),提出AgenticInterleave框架和IVR-12评分标准,发现前沿VLM准确率最高52.8%,交错答案质量与人工参考差距显著。

AI中文摘要:

视觉语言模型(VLM)日益驱动面向消费者的AI搜索,但在日常视觉问题的多样性上对其进行评估仍然具有挑战性。现有基准通常针对预定义能力,如多跳检索或长形式合成,而用户提出的基于照片的问题涵盖了日常场景的长尾分布。尽管VLM取得了进步,但在主流中文图片分享平台小红书上,用户仍会向他人寻求日常视觉问题的帮助。受此行为启发,我们从这些问题中筛选构建了NoteVQA,共包含12个主题类别和7种用户意图下的252个条目。每个条目包含一个从专家社区回复中提炼的简洁参考,以及一个经过人工审核的交错参考答案,该答案结合了文本解释与支持性视觉证据。我们评估了短答案正确性和交错答案质量。为支持后者,我们引入了AgenticInterleave,一个用于检索支持答案生成的单智能体ReAct框架,以及IVR-12,一个用于评估交错参考和模型输出的内容、呈现和图像质量的12维评分标准。在10个前沿VLM中,最高短答案准确率为52.8%,而为Qwen3.5-397B-A17B添加智能体搜索仅将准确率提高了2.0%。对于交错答案,运行AgenticInterleave的同一模型在IVR-12下得分为3.52,而人工审核参考为4.65,其中内容质量差距最大。这些结果凸显了日常视觉问题对当前VLM在答案准确性和视觉基础解释质量方面构成的挑战。

英文摘要:

Vision-language models (VLMs) increasingly power consumer-facing AI search, yet evaluating them on the diversity of everyday visual questions remains challenging. Existing benchmarks often target predefined capabilities, such as multi-hop retrieval or long-form synthesis, whereas users ask photo-grounded questions spanning a long tail of everyday scenarios. Despite advances in VLMs, users on Xiaohongshu, a mainstream Chinese image-sharing platform, continue to turn to other people for help with everyday visual questions. Motivated by this behaviour, we curate NoteVQA from these questions, yielding 252 items across 12 topical categories and 7 user intents. Each item includes a concise reference distilled from expert community responses and a human-audited interleaved reference answer that combines textual explanations with supporting visual evidence. We evaluate both short-answer correctness and interleaved-answer quality. To support the latter, we introduce AgenticInterleave, a single-agent ReAct framework for retrieval-supported answer generation, together with IVR-12, a 12-dimensional rubric for assessing the content, presentation, and image quality of interleaved references and model outputs. Across 9 frontier VLMs, the highest short-answer accuracy is 52.8\%, while adding agentic search to Qwen3.5-397B-A17B improves accuracy by only 2.0\%. For interleaved answers, the same model running AgenticInterleave scores 3.52 under IVR-12, compared with 4.65 for the human-audited references, with the largest gap in content quality. These results highlight the challenges that everyday visual questions pose for current VLMs in both answer accuracy and the quality of visually grounded explanations.

补充信息

↑