发表机构
University of California, Santa Barbara(加州大学圣巴巴拉分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明多模态情感识别中 oracle 互补性上限并非融合目标,实际增益受编码器来源影响,拼接优于学习路由器,提升空间反映分支分歧而非融合预算。
AI 中文摘要
多模态系统的互补性分析通常报告一个 oracle 上限(至少一个单模态分支正确的样本比例),并将该上限与实现的融合准确率之间的差距解释为可恢复的提升空间。我们表明,该上限并非融合目标。使用冻结的自监督音频和视觉编码器及训练好的头部,我们在 CREMA-D 数据集上测量了两种音频编码器(一种的微调谱系包含 CREMA-D,另一种为干净的编码器)下的上限和实际增益,并在 EAV 数据集上使用受污染的编码器进行测量。将微调编码器替换为自监督编码器使提升空间增加了两倍,从 +0.055 增至 +0.186。第二个干净编码器和对受污染编码器的受控退化落在同一条曲线上,因此来源通过改变音频分支的准确率来移动提升空间。拼接最多实现约一半的提升空间(转换分数 0.55),而在 EAV 上,尽管提升空间更大,转换分数与零无法区分。在所有情况下,学习路由器均被普通拼接击败。Oracle 提升空间衡量的是分支分歧,而非融合预算。
英文摘要
Complementarity analyses of multimodal systems commonly report an oracle ceiling (the fraction of examples on which at least one unimodal branch is correct) and interpret the gap between it and realized fusion accuracy as recoverable headroom. We show that this ceiling is not a fusion target. Using frozen self-supervised audio and visual encoders with a trained head, we measure the ceiling and realized gain on CREMA-D under two audio encoders (one whose fine-tuning lineage includes CREMA-D, one clean) and on EAV under the contaminated encoder. Replacing the fine-tuned encoder with a self-supervised one triples the headroom, from +0.055to +0.186. A second clean encoder and a controlled degradation of the contaminated one fall on the same curve, so provenance moves the headroom by changing audio-branch accuracy. Concatenation realizes at most about half of the headroom (converted fraction 0.55), and on EAV the converted fraction is indistinguishable from zero despite larger headroom. A learned router is beaten by plain concatenation everywhere. Oracle headroom measures branch disagreement, not a fusion budget.
Comments4+1 pages, 1 figure, 1 table. Submitted to ICASSP 2027