AI 中文总结
提出PRIME闭环可靠性引导框架,通过诊断恢复多模态意图识别的不可靠模态,在保持干净数据性能的同时提升了多模态数据各类退化场景下的鲁棒性。
AI 中文摘要
多模态意图识别结合了语言、声学和视觉证据,但各模态可能存在噪声、缺失、语义冲突或不成比例地占主导地位的问题。现有方法通常隐式推断模态重要性,对不可靠输入进行重加权或抑制,却未确定退化的模态是否可以修复并随后被信任。我们提出PRIME(Precision-weighted Reliability Inference and Modality rEstoration,即精度加权可靠性推理与模态恢复),这是一个闭环可靠性引导框架,在样本级别联合诊断、恢复和重新评估模态质量。PRIME通过从互补诊断证据(包括预测置信度、认知分歧、跨模态共识和特征退化)中估计的上下文对数方差来表示每个模态的弱点。由于模态可靠性标注不可用,估计器使用具有已知退化程度的受控模态损坏,结合异方差不确定性目标进行显式训练。PRIME不直接丢弃不可靠模态,而是利用其估计的弱点控制原型条件变分恢复模块,从互补模态重建退化表示。关键在于,恢复后会重新估计可靠性,使模型能够确定修复后的表示是否足够可信以用于预测。恢复后的精度用于逆方差多模态融合。在多模态意图识别基准上的实验表明,PRIME在保持有竞争力的干净数据性能的同时,提高了在缺失、噪声、冲突和模态不平衡条件下的鲁棒性。
英文摘要
Multimodal intent recognition combines linguistic, acoustic, and visual evidence, but individual modalities may be noisy, missing, semantically conflicting, or disproportionately dominant. Existing methods typically infer modality importance implicitly and either reweight or suppress unreliable inputs, without determining whether a degraded modality can be repaired and subsequently trusted. We propose PRIME (Precision-weighted Reliability Inference and Modality rEstoration), a closed-loop reliability guided framework that jointly diagnoses, restores, and reassesses modality quality at the sample level. PRIME represents the weakness of each modality through a contextual log-variance estimated from complementary diagnostic evidence, including predictive confidence, epistemic disagreement, cross-modal consensus, and feature degeneracy. Because modality-reliability annotations are unavailable, the estimator is explicitly trained using controlled modality corruption with known degradation severity, together with a heteroscedastic uncertainty objective. Rather than directly discarding an unreliable modality, PRIME uses its estimated weakness to control a prototype-conditioned variational restoration module that reconstructs the degraded representation from complementary modalities. Crucially, reliability is re-estimated after restoration, allowing the model to determine whether the repaired representation has become sufficiently trustworthy to contribute to prediction. The resulting post-restoration precisions are used for inverse-variance multimodal fusion. Experiments on multimodal intent-recognition benchmarks show that PRIME maintains competitive clean-data performance while improving robustness under missing, noisy, conflicting, and modality-imbalanced conditions.