arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态智能体搜索中的无声故障:一种诊断分类法与跨评判者评估

Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation

Zhengxian Wu, Junjie Gao, Kai Yang

arXiv 2607.19793首次发表:更新:

发表机构

Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多模态智能体搜索中无声故障这一隐藏可靠性问题,引入六类分类法,构建轨迹级诊断管道,通过实验表明表面准确性高估真实正确性,且无声故障与能力相关、常转移。

AI 中文摘要

多模态智能体搜索系统越来越依赖外部工具来回答知识密集型视觉问题。然而,现有评估主要关注最终答案的准确性,可能会忽略搜索轨迹中的故障。在这项工作中,我们研究了诸如无声故障等隐藏的可靠性问题。我们引入了一个六类分类法,涵盖模态捷径、幻影接地、错误证据正确答案情况、过度检索清洗、跨模态矛盾和来源幻觉。基于此分类法,我们构建了一个轨迹级诊断管道,在统一的ReAct风格框架下评估答案正确性和证据接地质量。对四个前沿多模态模型的MMSearch-Plus轨迹进行的实验表明,表面准确性一直高估了真实轨迹级的正确性。我们进一步使用跨评判者验证、空白图像压力测试和工具消融来表明,无声故障与能力相关,并且通常会转移而不是消失。

英文摘要

Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions. However, existing evaluations mainly focus on final-answer accuracy and may miss failures in the search trajectory. In this work, we study such hidden reliability issues as silent failures. We introduce a six-category taxonomy covering modality shortcuts, phantom grounding, wrong-evidence-right-answer cases, over-retrieval laundering, cross-modal contradiction, and provenance hallucination. Based on this taxonomy, we build a trajectory-level diagnostic pipeline that evaluates both answer correctness and evidence-grounding quality under a unified ReAct-style scaffold. Experiments on MMSearch-Plus trajectories across four frontier multimodal models show that surface accuracy consistently overestimates true trajectory-level correctness. We further use cross-judge validation, blank-image stress tests, and tool ablations to show that silent failures are capability-dependent and often shift rather than disappear. Home-page: https://github.com/DingWu1021/silent-failures-multimodal-agentic-search

Journal refSIGIR 2026 SynthIR Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑