arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37691eess.AS

信号无关与信号相关的任意阵列可变麦克风数量神经Ambisonic矩阵编码

Signal-Independent and Signal-Dependent Neural Ambisonic Matrix Encoding for Arbitrary Arrays with Variable Microphone Counts

  • Soochow University(苏州大学)

机构由 AI 辅助整理,请以论文原文为准。

Shichao Hu, Zhiheng Jin, Chunyang Xu, Mengyao Zhu

中文总结 AI 辅助

针对固定麦克风数量限制,提出基于Transformer的任意阵列可变麦克风数量Ambisonic矩阵编码,含信号无关与信号相关扩展,优于传统最小二乘编码,实现阵列无关且泛化良好。

中文摘要 AI 辅助

最近的神经Ambisonic编码器能够适应多样的阵列几何形状,然而许多现有的神经编码器需要固定的麦克风数量,因为麦克风通道的数量嵌入在网络架构中。这一要求限制了在不同麦克风配置的设备上的部署,以及对可用通道变化的适应性。为了解决这一限制,我们研究了基于Transformer的矩阵编码,用于具有可变麦克风数量的任意麦克风阵列。这是通过共享的逐麦克风处理和掩蔽自注意力来实现的,该注意力对可变大小阵列中的麦克风间关系进行建模。在此框架内,我们考虑了信号无关(SI)编码,它从阵列传递函数预测编码矩阵,并引入了信号相关(SD)扩展,该扩展额外结合了观测到的麦克风信号。两种模型都在使用LibriSpeech源的模拟场景上进行了训练,并在源类型变化、未见过的麦克风数量以及训练期间使用的源数量增加的情况下进行了广泛评估。SI和SD在评估条件下的总体重建性能上都优于传统的(LS)编码。SD始终比SI获得更强的整体性能。这些结果表明,所提出的框架能够实现阵列无关的Ambisonic编码,同时保持跨麦克风数量和声源条件的泛化能力。

英文摘要

Recent neural Ambisonic encoders accommodate diverse array geometries, yet many existing neural encoders require a fixed microphone count because the number of microphone channels is embedded in the network architecture. This requirement limits deployment across devices with different microphone configurations and adaptation to changes in available channels. To address this limitation, we investigate Transformer-based matrix encoding for arbitrary microphone arrays with variable microphone counts. This is achieved through shared microphone-wise processing and masked self-attention that models inter-microphone relationships across variable-size arrays. Within this framework, we consider signal-independent (SI) encoding, which predicts encoding matrices from array transfer functions, and introduce a signal-dependent (SD) extension that additionally incorporates the observed microphone signals. Both models are trained on simulated scenes using LibriSpeech sources and extensively evaluated under changes in source type, unseen microphone counts, and increased source counts beyond those used during training. Both SI and SD outperform conventional least-squares (LS) encoding in aggregate reconstruction performance across the evaluated conditions. SD consistently achieves stronger overall performance than SI. These results demonstrate that the proposed framework enables array-agnostic Ambisonic encoding while retaining generalization across microphone counts and acoustic source conditions.

↑