通过可学习观测前端进行语义采样
Semantic Sampling via Learnable Observation Front Ends
浏览论文内容
中文总结 AI 辅助
研究针对声学信号重建,提出通过可学习观测前端进行语义采样的方法,该前端含语义特征滤波器组等组件,经实验验证,在相同观测预算下比其他方法能提供更多信息观测,为声学信号重建保留更多有用信息。
中文摘要 AI 辅助
采样决定了下游重建系统可用信息的形式。传统低速率采样直接从原始波形生成有限维观测,采样规则主要受带宽、稀疏性或固定信号级结构引导。对于语音等声学信号,与重建相关的信息常通过与内容相关的谱-时结构而非仅波形样本表达。本文提出通过可学习观测前端进行语义采样,从学习到的信号响应生成有限维观测而非直接对波形点进行下采样。所提出的前端由语义特征滤波器组、受限语义观测矩阵和低速率读出模块组成。滤波器组将输入波形映射到多个声学响应通道,观测矩阵将这些响应组合成少量观测通道,读出模块产生低速率有限维样本。然后使用重建网络从所得观测中恢复信号。低速率语音重建实验表明,在相同观测预算下,所提出的语义采样前端比基于预定低速率波形的固定低速率采样和神经恢复方法提供更多信息观测。波形保真度、谱一致性和感知质量的提升表明,在相同观测预算下,可学习观测前端为声学信号重建保留了更多有用信息。
英文摘要
Sampling determines the form of information available to downstream reconstruction systems. Conventional lowrate sampling forms finite-dimensional observations directly from the raw waveform, with the sampling rule mainly guided by bandwidth, sparsity, or fixed signal-level structures. For acoustic signals such as speech, however, reconstruction-relevant information is often expressed through content-related spectral-temporal structures rather than waveform samples alone. This paper proposes semantic sampling via learnable observation front ends, where finite-dimensional observations are generated from learned signal responses instead of directly subsampled waveform points. The proposed front end consists of a semantic feature filterbank, a constrained semantic observation matrix, and a low-rate readout module. The filterbank maps the input waveform into multiple acoustic response channels, the observation matrix combines these responses into a small number of observation channels, and the readout module produces low-rate finite-dimensional samples. A reconstruction network is then used to recover the signal from the resulting observations. Experiments on low-rate speech reconstruction show that, under the same observation budget, the proposed semantic sampling front end provides more informative observations than fixed low-rate sampling and neural restoration methods based on predetermined low-rate waveforms. The improvements in waveform fidelity, spectral consistency, and perceptual quality show that learnable observation front ends preserve more useful information for acoustic signal reconstruction under the same observation budget.