双BEATs:通过抖动在音频大语言模型中解锁零样本立体声音频感知
Dual-BEATs: Unlocking Zero-Shot Stereo Audio Perception in Audio Large Language Models via Dithering
- Institute of Information Science, Academia Sinica(台湾中央研究院资讯科学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对多模态大语言模型空间感知局限,提出双BEATs架构,通过在编码前注入抖动噪声解决归一化问题,在三元方向分类任务中验证该方法有出色空间分辨率且能零样本泛化,证明标准模型经正则化可实现广义立体声音频理解。
AI中文摘要:
多模态大语言模型有出色的语义音频理解能力,但因依赖单声道音频表示而在空间感知上存在局限。当前空间音频感知方法受限。本文引入双BEATs架构,左右声道独立通过相同语义编码器。为避免归一化消除立体声通道间差异,编码前注入静态不相关抖动噪声。在三元方向分类任务中评估,抖动模型有出色空间分辨率,能零样本泛化到未见空间配置,表明适当正则化下标准多模态模型能实现广义立体声音频理解。
英文摘要:
Multimodal Large Language Models (LLMs) have remarkable semantic audio understanding, yet they remain "spatially agnostic" due to their reliance on mono-channel audio representations. Currently, spatial audio perception methods mainly focus on complex room simulations and custom-trained, geometry-aware stereo encoders, which limits their accessibility and generalizability. In this paper, we introduce the Dual-BEATs architecture, in which the left and right audio channels are routed independently through two identical semantic encoders as an alternative to specialized spatial modules. To circumvent the architectural bottleneck where internal normalization otherwise erases the inter-channel variance of stereo audio, we inject a static, uncorrelated dithering noise floor prior to encoding. This dithering intervention establishes a macro-variance floor that "smuggles" spatial geometry across the normalization layers. Evaluated on a ternary directional classification task (Left, Center, Right), we demonstrate that dithered models achieve exceptional spatial resolution--reaching up to 97.2% localization accuracy even on subtle 0.5 panning amplitudes--and demonstrates robust, zero-shot generalization to entirely unseen spatial configurations. Our results suggest that with the appropriate acoustic regularization, standard multimodal models are natively capable of generalized stereo audio understanding.