arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向以房间冲激响应(RIR)表征的声学环境的深度神经压缩:结构感知约束

Deep Neural Compression for RIR-Characterized Acoustic Environments with Structure-Aware Constraints

Chen-Yuan Ning, Yang Ai, Hui-Peng Du, Xiao-Hang Jiang, Zhen-Hua Ling

arXiv 2609.04085首次发表:更新:

发表机构

National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China(语音和语言信息处理国家重点实验室,中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对RIR存储负担问题,提出基于EnCodec的神经RIR压缩方法,引入结构感知约束,在低比特率下实现更优的RIR重建与混响语音感知性能。

AI 中文摘要

房间冲激响应(RIR)通过捕捉声音在封闭空间内的传播与衰减特性来表征房间的声学环境。在沉浸式音频渲染等应用中,精确的声学重建往往依赖于空间密集采样的RIR,这会产生大量RIR数据,给存储带来沉重负担。尽管近期的神经音频编解码器为低比特率压缩提供了有效框架,但其训练目标主要针对语音和通用音频,与RIR的声学特性并不匹配。因此,本文提出一种基于EnCodec的神经RIR压缩方法,在两个层面引入RIR结构感知约束。具体而言,在RIR层面,通过能量衰减曲线(EDC)正则化和短时窗能量约束,对RIR的全局衰减行为与局部能量分布施加结构感知约束;在混响语音层面,进一步引入混响语音监督,以约束由重建RIR生成的混响语音的一致性。实验结果表明,在375 bps的低比特率下,该方法相比面向音频的编解码器,实现了更低的RIR重建误差和更优的混响语音感知一致性。

英文摘要

Room impulse responses (RIRs) characterize the acoustic environment of a room by capturing how sound propagates and decays within an enclosed space. In applications such as immersive audio rendering, accurate acoustic reconstruction often relies on spatially densely sampled RIRs. This consequently gives rise to a large volume of RIR data, imposing a substantial burden on storage. Although recent neural audio codecs provide an effective framework for low-bitrate compression, their training objectives are mainly tailored to speech and general audio, and are therefore not well aligned with the acoustic characteristics of RIRs. Therefore, we propose an EnCodec-based neural RIR compression method, which incorporates RIR structure-aware constraints at two levels. Specifically, at the RIR level, structure-aware constraints are imposed on the global decay behavior and local energy distribution of RIRs through energy decay curve (EDC) regularization and a short-time window energy constraint, while at the reverberant-speech level, reverberant-speech supervision is further introduced to constrain the consistency of the reverberant speech generated by the reconstructed RIRs. Experimental results show that, at a low bitrate of 375 bps, the proposed method achieves lower RIR reconstruction error and better reverberant-speech perceptual consistency than audio-oriented codecs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑