发表机构
International Institute of Information Technology Bangalore(国际信息技术学院班加罗尔)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出残差全速率PR架构,将AddUNet等价于多速率PR滤波器组,通过结构性守恒实现任务导向路由,在TIMIT上显著降低PER并保持精确重建。
AI 中文摘要
本文建立了AddUNet的完美重建(PR)解释及其全速率实现,并提出了一种用于任务导向表示学习的残差全速率PR架构。受约束的加性U-Net的幸存者-跳过结构被证明恰好等价于临界采样的多速率PR滤波器组。全速率公式在保持PR的同时,消除了临界采样系统的互补子带限制。随后提出了一种残差全速率PR架构,以逐步将任务无关、干扰或冗余结构从面向任务的幸存者中路由出去,同时显式保留被路由的信息。对于任意形状兼容的线性或非线性路由算子,精确重建均有保证,无需可逆性、匹配的综合滤波器组、重建损失或学习解码器。由此产生的架构将表示设计与重建设计解耦:守恒是结构性的,而学习则致力于任务导向的路由。相同的公式将具有恒等捷径的ResNet及其保留的残差输出识别为全速率PR系统。实验验证了线性可分离因子在机器精度下的精确单通道路由。在TIMIT上,所提出的前端在保持识别器和训练协议不变的情况下,将测试音素错误率(PER)从$28.60\pm2.09\\%$改善至$25.76\pm0.41\\%$,同时保持精确重建。说话人探测进一步表明,结构守恒本身并不隐含任务特定的不变性。
英文摘要
This paper establishes a perfect-reconstruction (PR) interpretation of AddUNet and its full-rate realization, and introduces a Residual Full-Rate PR architecture for task-directed representation learning. The survivor--skip structure of a constrained additive U-Net is shown to be exactly equivalent to a critically sampled multirate PR filter bank. The full-rate formulation removes the complementary-subband restrictions of the critically sampled system while preserving PR. A Residual Full-Rate PR architecture is then proposed to progressively route task-irrelevant, nuisance, or redundant structure away from the task-facing survivor while retaining the routed information explicitly. Exact reconstruction is guaranteed for arbitrary shape-compatible linear or nonlinear routing operators, without requiring invertibility, a matched synthesis bank, reconstruction loss, or learned decoder. The resulting architecture decouples representation design from reconstruction design: conservation is structural, while learning is devoted to task-directed routing. The same formulation identifies an identity-shortcut ResNet with its residual output retained as a full-rate PR system. Experiments verify exact single-channel routing of linearly separable factors to machine precision. On TIMIT, the proposed front-end improves test PER from $28.60\pm2.09\%$ to $25.76\pm0.41\%$ with the recognizer and training protocol held fixed, while maintaining exact reconstruction. Speaker probing further shows that structural conservation does not itself imply task-specific invariance.