arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态大语言模型为何失败及如何失败:基于因果任务分解的能力缺陷诊断

Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis

Xia Hu, Brian Potetz, Chun-Ta Lu, Huanfen Yao, Leonidas Guibas, Zhicheng Wang, Howard Zhou, Pengfei Xing, Andrew Gallagher

arXiv 2609.38851首次发表:更新:

发表机构

Google Research; Stanford University(谷歌研究院; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出因果分解框架CADET,通过受控干预前置依赖诊断MLLM失败源于内在缺陷还是级联错误,实验发现提供正确前置可消除认知任务54%错误,单一关键前置贡献占全部收益的84%。

AI 中文摘要

组合任务上的端到端准确率记录了多模态大语言模型(MLLMs)失败的频率,但无法区分一次失败是反映了目标能力的内在缺陷,还是来自上游前置条件的级联错误。我们提出一个因果分解框架,通过对每个任务的前置依赖关系进行受控干预,将这两种失败模式分离开来。我们的能力指标(NC、IC、RC)在无辅助、正确或错误的前置条件下对每个任务进行评分,以诊断失败产生的位置;贡献指标(N-Score、S-Score)改编自因果概率,量化每个前置条件的必要性和充分性,以确定失败的原因。我们将该框架实例化为CADET,一个诊断基准,包含10个复合任务,分解为46个单元任务,涵盖超过33,000个人工标注问题,涉及感知、空间、时间和认知类别。使用我们的框架诊断前沿MLLMs,揭示了端到端准确率所掩盖的系统性模式。在能力方面,提供正确的前置条件消除了认知任务上54%的错误,将其从最弱提升到优于空间和时间任务。在前置条件方面,因果贡献集中在少数关键前置条件上,仅提供最重要的单一前置条件就能捕获提供所有前置条件所获收益的84%。

英文摘要

End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through controlled interventions on the prerequisite dependencies of each task. Our capability metrics (NC, IC, RC) score each task under unassisted, correct, or incorrect prerequisites to diagnose where failures arise; contribution metrics (N-Score, S-Score), adapted from probabilities of causation, quantify each prerequisite's necessity and sufficiency to determine why. We instantiate the framework in CADET, a diagnostic benchmark of 10 composite tasks decomposed into 46 unit tasks with over 33,000 human-annotated questions spanning perception, spatial, temporal, and cognitive categories. Diagnosing frontier MLLMs with our framework uncovers systematic patterns that end-to-end accuracy obscures. Capability-wise, supplying correct prerequisites eliminates 54\% of errors on cognitive tasks, lifting them from weakest to above spatial and temporal. Prerequisite-wise, causal contributions are concentrated in a few critical prerequisites, and supplying the single most important one alone captures 84\% of the gain from supplying all prerequisites.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑