GAUGE:面向不完整多模态分类的粒度自适应反事实证据门控
GAUGE: Granularity-Adaptive Counterfactual Gating of Evidence for Incomplete Multimodal Classification
浏览论文内容
中文总结 AI 辅助
针对不完整多模态分类问题,提出轻量型GAUGE框架,通过细粒度反事实证据门控实现更优性能,为模态不完整下的证据控制提供了可扩展方案。
中文摘要 AI 辅助
多模态分类通常假设所有模态都可用,但现实世界的输入往往是不完整的。插补和动态融合可缓解这种不完整性,但现有方法在粗糙的模态层面操作,因此无法在同一恢复模态内保留可靠组件同时抑制误导性组件,从而损害预测可靠性。为解决此问题,我们提出GAUGE,一种用于不完整多模态分类的轻量型反事实门控框架。GAUGE首先用冻结的插补器插补缺失模态,并将观测到的和恢复的输入统一编码为细粒度证据单元。GAUGE不明确干预每个单元,而是通过预测感知泰勒证据分数来计算用参考表示替换每个单元的反事实效应分数,所有这些都在单次前向-后向传递中获得。这些分数被映射到连续门,转换为 additive attention-logit 偏置以进行单元级证据调制,且不改变骨干架构。在六个基准上的实验表明,GAUGE在各种不完整输入设置中均优于强大的基线。此外,泰勒余项理论分析表征了一阶近似相对于精确反事实效应的误差,确立GAUGE为模态不完整下细粒度证据控制的有原则且可扩展的框架。
英文摘要
Multimodal classification typically assumes all modalities are available, yet real-world inputs are often incomplete. Imputation and dynamic fusion can mitigate such incompleteness, but existing methods operate at a coarse modality level and thus cannot retain reliable components while suppressing misleading ones within the same recovered modality, compromising prediction reliability. To address this issue, we propose GAUGE, a lightweight counterfactual gating framework for incomplete multimodal classification. GAUGE first imputes missing modalities with a frozen imputer and encodes observed and recovered inputs uniformly as fine-grained evidence units. Rather than intervening on each unit explicitly, GAUGE scores the counterfactual effect of replacing every unit with a reference representation through prediction-aware Taylor evidence scores, all obtained in a single forward-backward pass. These scores are mapped to continuous gates, which are converted into additive attention-logit biases for unit-wise evidence modulation without altering the backbone architecture. Experiments across six benchmarks demonstrate that GAUGE outperforms strong baselines across diverse incomplete-input settings. Furthermore, a Taylor remainder theoretical analysis characterizes the error of the first-order approximation relative to the exact counterfactual effect, establishing GAUGE as a principled and scalable framework for fine-grained evidence control under modality incompleteness.