多声道音频质量评估:将预训练感知模型扩展到空间音频
Multichannel Audio Quality Assessment: Extending Pretrained Perceptual Models to Spatial Audio
浏览论文内容
中文总结 AI 辅助
针对多声道空间音频质量评估,提出特征级FGAtt方法,在5.1声道测试集上表现最佳,有效复用预训练感知模型。
中文摘要 AI 辅助
准确的感知质量评估对于评估和优化空间音频至关重要,其中感知质量取决于信号保真度和声道间空间关系。然而,主观评估成本高昂,而现有的感知模型通常针对有限的声道配置进行训练,无法直接应用于更高声道数的音频。这引出了一个问题:如何有效复用预训练的感知知识来处理多声道空间音频?我们使用5.1声道音频,研究了四种多声道集成级别:信号级、预测级、潜在级和特征级,并提出了两种学习方法:空间组表示的潜在级聚合,以及特征级特征带组注意力(FGAtt),后者在感知处理之前自适应地融合特征级别的空间组。在五个5.1声道测试集上,FGAtt取得了最强的整体性能,证明了特征级自适应对于复用预训练感知知识的有效性。
英文摘要
Accurate perceptual quality assessment is essential for evaluating and optimizing spatial audio, where perceived quality depends on both signal fidelity and inter-channel spatial relationships. However, subjective evaluation is costly, while existing perceptual models are often trained for limited channel configurations and cannot be directly applied to higher-channel-count audio. This raises the question: how can pretrained perceptual knowledge be effectively reused for multichannel spatial audio? Using 5.1-channel audio, we study four levels of multichannel integration: signal, prediction, latent, and feature and propose two learned approaches: latent-level aggregation of spatial-group representations and the feature-level Feature-Band Group Attention (FGAtt), which adaptively fuses spatial groups at the feature level before perceptual processing. Across five 5.1-channel test sets, FGAtt achieves the strongest over- all performance, demonstrating the effectiveness of feature-level adaptation for reusing pretrained perceptual knowledge
发表机构
- Dolby Laboratories(杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。