感知启发的贝叶斯因果融合用于视听源定位
Perception-Inspired Bayesian Causal Fusion for Audiovisual Source Localization
浏览论文内容
中文总结 AI 辅助
提出一种受人类感知启发的贝叶斯因果融合方法,通过推断共同原因后验来门控视听融合,避免无关模态干扰,在无需重训练下提升定位精度并降低误差。
中文摘要 AI 辅助
多模态融合有望实现更准确的感知,但前提是各模态共享一个共同原因。当它们不共享时,第二模态不携带关于目标的信息,融合它只会破坏估计。我们将这种是否融合的决策视为贝叶斯因果推断,遵循人类多感官感知的最优观察者模型,并将其实现为冻结的音频和视觉模型之上的即插即用层,用于声音事件定位和检测。该模型推断可见候选上的共同原因后验,然后相应地门控精度加权融合。无条件融合会使方向误差增加一倍以上,而因果门控在限制屏幕外退化的同时改善了屏幕上的定位,且无需任何联合网络重新训练。
英文摘要
Multimodal fusion promises more accurate perception but only when the modalities share a common cause. When they do not, the second modality carries no information about the target, and fusing it can only corrupt the estimate. We cast this whether-to-fuse decision as Bayesian causal inference, following the optimal-observer model of human multisensory perception, and implement it as a plug-and-play layer on top of frozen audio and visual models for sound event localization and detection. The model infers a common-cause posterior over visible candidates, then gates precision-weighted fusion accordingly. Fusing unconditionally more than doubles the direction error, whereas the causal gate improves on-screen localization while limiting off-screen degradation, without any joint network retraining.
发表机构
- Aalto University(阿尔托大学)
- Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
- Friedrich-Alexander-Universität Erlangen-Nürnberg(埃尔朗根-纽伦堡大学)
机构由 AI 辅助整理,请以论文原文为准。