C³PO:评估全模态模型中的跨模态组合与反事实性能
C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models
- Microsoft Research India(微软研究院印度分院)
- IIIT Hyderabad(印度国际信息技术研究院(海得拉巴))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究推出C³PO基准数据集,评估全模态模型的跨模态组合与反事实性能,发现模型存在模态主导问题,需改进架构以支持持续跨模态注意力。
AI中文摘要:
当前多模态大语言模型(MLLMs)可处理多种感官输入,但其推理仍严重偏向主导模态,导致跨模态推理脆弱。我们推出C³PO,这一包含3404个样本的基准数据集,涵盖视频、音频、图像和文本,评估两项能力:信息组合(融合分散证据)与反事实冲突(解决刻意矛盾)。C³PO的配对IC/CC结构与四层设计,可针对性诊断跨模态推理失败的时机与原因。通过使用25个逻辑严谨模板的全自动流程构建,C³PO显示人类准确率达88.64%,而最优模型Gemini-3.1-Pro仅达73.17%,开源模型在冲突下表现崩溃。通过注意力探测,我们发现86%-95%的失败源于模态主导性:模型偏向单一模态而忽略矛盾证据,87%-95%的注意力集中在文本上。中层注意力熵可预测正确性——持续探索成功,过早崩溃失败。同等复杂模板间56点的准确率差距表明,性能取决于模态在冲突解决中的结构角色,而非组合方式。这些发现显示多模态感知不保证稳健推理,架构必须支持持续跨模态注意力以避免过早
英文摘要:
Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning. We introduce C$^3$PO, a benchmark of 3,404 samples spanning video, audio, image, and text, evaluating two abilities: information composition (fusing dispersed evidence) and counterfactual conflict (resolving deliberate contradictions). C$^3$PO's paired IC/CC structure and four-tier design enable targeted diagnosis of when and why cross-modal reasoning fails. Built through a fully automatic pipeline using 25 logically grounded templates, C$^3$PO reveals that while humans achieve 88.64% accuracy, the best model (Gemini-3.1-Pro) reaches only 73.17%, with open-source models collapsing under conflict. Through attention probes, we find 86-95% of failures stem from modality dominance: models commit to one modality while ignoring contradictory evidence, concentrating 87-95% of attention on text. Mid-layer attention entropy predicts correctness-sustained exploration succeeds, premature collapse fails. The 56-point accuracy gap between equally complex templates reveals that performance depends on modalities' structural roles in conflict resolution, not combinations. These findings show multimodal perception does not guarantee robust reasoning; architectures must enable sustained cross-modal attention to avoid premature