arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缺失模态下的共形融合

Conformal Fusion Under Missing Modalities

Alireza Moayedikia

arXiv 2608.07183首次发表:更新:

发表机构

Swinburne University of Technology(斯威本科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MCCF架构,解决缺失模态下的多模态融合问题,该架构兼具模态缺失鲁棒性与校准不确定性,在多基准测试中表现出良好的覆盖率与精度性能。

AI 中文摘要

多模态融合架构通常假设推理时所有模态均可用,但传感器故障、采集变异性和成本约束常会导致观测不完整。现有研究将模态缺失视为预测精度问题,却未解答一个更基础的问题:当整个输入流被移除时,模型的置信度估计是否仍保持校准。本文提出模态条件共形融合(Modality-Conditioned Conformal Fusion, MCCF),该架构同时解决模态缺失鲁棒性与校准不确定性问题。MCCF结合经模态丢弃训练的多模态瓶颈融合主干、生成模态分解狄利克雷分布的单模态证据头,以及将单模态证据融合为联合预测分布的Dempster-Shafer组合规则;缺失模态贡献空证据,会被结构性忽略,因此融合后的不确定性自动反映信息减少,无需测试时插补。基于模态存在掩码的Mondrian共形校准模块,为每个非空模态子集提供有限样本组条件覆盖率。据本文所知,MCCF是首个通过架构整合而非事后重新校准,在任意模态可用性下具有形式化覆盖率保证的方法,且证据分解产生的单模态空值分数可将不确定性定位到导致缺失的模态。在合成问题和三个真实多模态基准上,MCCF在所有模态存在子集上保持目标覆盖率,相比边际分裂共形基线大幅缩小全模态与部分模态间的覆盖率差距,且相对于温度缩放和证据基线无明显精度损失。

英文摘要

Multimodal fusion architectures typically assume all modalities are available at inference, yet sensor failures, acquisition variability, and cost constraints routinely produce incomplete observations. Existing work treats modality absence as a prediction-accuracy problem, leaving a more basic question unanswered: whether a model's confidence estimates remain calibrated when an entire input stream is removed. We argue that missing-modality robustness and calibrated uncertainty are a single coupled property, and introduce Modality-Conditioned Conformal Fusion (MCCF), an architecture that addresses both at once. MCCF combines a multimodal bottleneck fusion backbone trained with modality dropout, per-modality evidential heads producing modality-decomposed Dirichlet distributions, and a Dempster-Shafer combination rule that fuses the per-modality evidence into a joint predictive distribution; an absent modality contributes vacuous evidence that is structurally ignored, so the fused uncertainty automatically reflects the reduced information without test-time imputation. A Mondrian conformal calibration module keyed on the modality-presence mask then provides finite-sample group-conditional coverage for every non-empty modality subset. MCCF is, to our knowledge, the first method with formal coverage guarantees under arbitrary modality availability through architectural integration rather than post-hoc recalibration, and the evidential decomposition yields per-modality vacuity scores that localise uncertainty to the absent modality responsible. Across a synthetic problem and three real multimodal benchmarks, MCCF holds its target coverage on every modality-presence subset, substantially narrows the coverage gap between full and partial modalities relative to a marginal split-conformal baseline, and imposes no measurable accuracy cost relative to temperature-scaled and evidential baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑