arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36441eess.AS

感知启发的贝叶斯因果融合用于视听源定位

Perception-Inspired Bayesian Causal Fusion for Audiovisual Source Localization

Kyung Yun Lee, Sungnyun Kim, Sebastian J. Schlecht, Tae-Hyun Oh, Vesa Välimäki

首次发表
浏览论文内容

中文总结 AI 辅助

提出一种受人类感知启发的贝叶斯因果融合方法,通过推断共同原因后验来门控视听融合,避免无关模态干扰,在无需重训练下提升定位精度并降低误差。

中文摘要 AI 辅助

多模态融合有望实现更准确的感知,但前提是各模态共享一个共同原因。当它们不共享时,第二模态不携带关于目标的信息,融合它只会破坏估计。我们将这种是否融合的决策视为贝叶斯因果推断,遵循人类多感官感知的最优观察者模型,并将其实现为冻结的音频和视觉模型之上的即插即用层,用于声音事件定位和检测。该模型推断可见候选上的共同原因后验,然后相应地门控精度加权融合。无条件融合会使方向误差增加一倍以上,而因果门控在限制屏幕外退化的同时改善了屏幕上的定位,且无需任何联合网络重新训练。

英文摘要

Multimodal fusion promises more accurate perception but only when the modalities share a common cause. When they do not, the second modality carries no information about the target, and fusing it can only corrupt the estimate. We cast this whether-to-fuse decision as Bayesian causal inference, following the optimal-observer model of human multisensory perception, and implement it as a plug-and-play layer on top of frozen audio and visual models for sound event localization and detection. The model infers a common-cause posterior over visible candidates, then gates precision-weighted fusion accordingly. Fusing unconditionally more than doubles the direction error, whereas the causal gate improves on-screen localization while limiting off-screen degradation, without any joint network retraining.

发表机构

  • Aalto University(阿尔托大学)
  • Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
  • Friedrich-Alexander-Universität Erlangen-Nürnberg(埃尔朗根-纽伦堡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑