arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29278cs.CL

模态断层线:结构损坏揭示脆弱的全模态推理

Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning

Zhaolu Kang, Meixin Wu, Yu Xue, Yingjie He, Qiming Shi, Lei Wei, Yidi Wang, Richeng Xuan, Zhichao Hu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出SCEval诊断评估协议,通过结构损坏揭示全模态模型的脆弱推理,发现文本-视觉损坏是最稳定的共享断层线,多模态退化非相加性,干净准确率不代表结构不可靠时仍可靠。

中文摘要 AI 辅助

全模态大语言模型越来越多地在干净的文本-视觉-音频输入上接受评估,其中每个通道都存在、同步且易于解释。这类评估结果常被视为稳健跨模态融合的证据,但干净的评估无法判断成功是否依赖稳定的跨模态结构,还是仅依赖完整输入中足够的线索。为解决这一差距,我们定义了模态断层线:当某一模态仍存在且人类可解释,但其内部证据结构被扰动时,模型行为变得不稳定的边界。我们引入SCEval(结构损坏评估),这是一种诊断评估协议,在保持问题、答案空间和模态通道固定的同时,分别及联合对文本、视觉和音频应用受控的结构损坏。SCEval由来自Social-IQ、OmniBench和VALOR的273个人类验证的三模态示例构建,评估了15个专有和开源的全模态系统。结果表明,结构损坏会降低干净准确率,文本-视觉损坏形成最稳定的共享断层线,多模态退化是非相加性的,而非损坏模态数量的简单函数。因此,干净的全模态准确率并不能保证模型在跨模态证据变得结构不可靠时仍保持可靠。

英文摘要

Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and readily interpretable. Such scores are often taken as evidence of robust cross-modal fusion, but clean evaluation cannot tell whether success depends on stable cross-modal structure or on cues sufficient only in intact inputs. To address this gap, we define a modality fault line: a boundary at which model behavior becomes unstable when a modality remains present and human-interpretable, but its internal evidence structure is perturbed. We introduce SCEval (Structure-Corruption Evaluation) a diagnostic evaluation protocol that keeps the question, answer space, and modality channels fixed while applying controlled structural corruptions to text, vision, and audio individually and jointly. Built from $273$ human-verified tri-modal examples from Social-IQ, OmniBench, and VALOR, SCEval evaluates $15$ proprietary and open-source omni-modal systems. The results show that structural corruption lowers clean accuracy, text--vision damage forms the most stable shared fault line, and multi-modal degradation is non-additive rather than a simple function of the number of corrupted modalities. Clean omni-modal accuracy therefore does not establish that a model will remain reliable when cross-modal evidence becomes structurally unreliable.

发表机构

  • Tencent(腾讯)
  • Peking University(北京大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑