发表机构
Graduate School of Informatics and Engineering, the University of Electro-Communications; Department of Electrical Engineering and Computer Science, Tokyo University of Agriculture and Technology(东京电机大学信息与工程研究生院; 东京农业大学电气与计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出复值玻尔兹曼机PolarBM及其对数尺度变体LogPolarBM,能在极坐标和对数极坐标中处理音频信号,明确建模幅度和相位依赖关系,实验证明其比传统模型有更高建模精度,在多领域有广泛应用潜力。
AI 中文摘要
虽然大量数据(如音频信号频谱)自然地用复数表示,但传统机器学习方法常通过为实值变量设计的框架简化复域问题。本文提出一种新颖的玻尔兹曼机PolarBM,能在极坐标中自然处理复值变量,定义相位明确依赖于幅度的复变量概率密度函数。还提出LogPolarBM在对数尺度上对幅度建模,产生灵活的条件概率密度函数。实验表明,与传统模型相比,所提RBM通过明确建模幅度和相位之间的依赖关系实现了更高的建模精度。虽实验聚焦音频信号,但这些玻尔兹曼机的效用不限于音频应用,在涉及复值数据的多个科学和工程领域有广泛潜力。
英文摘要
Although vast amounts of data, such as audio signal spectra, are naturally represented using complex numbers, conventional machine learning methods often simplify complex-domain problems by employing frameworks designed for real-valued variables. While this simplification offers computational benefits, it discards structural information regarding the inherent relationship between amplitude and phase. In this paper, we propose a novel Boltzmann machine (BM), named PolarBM, capable of naturally handling complex-valued variables in the polar coordinate (i.e., an amplitude-phase representation). PolarBM defines a probability density function for complex variables in which the phase explicitly depends on the amplitude, thereby capturing the physically important relationships of complex-valued signals. Furthermore, to process audio signals in accordance with human auditory perception, we propose LogPolarBM, which models amplitude on a logarithmic scale. This extension yields a flexible conditional probability density function, a power-weighted noncentral complex Gaussian (PW-NCCG) distribution, whose marginal amplitude distribution encompasses the Rice, Nakagami, and noncentral chi distributions as special cases. For practical applications, we also introduce the restricted variants of these proposed models: PolarRBM and LogPolarRBM. Experimental results demonstrate that by explicitly modeling the dependency between amplitude and phase, the proposed RBMs achieve superior modeling accuracy compared to conventional models, including deep neural networks. Although our experiments focus on audio signals, the utility of the proposed BMs is not limited to audio applications; their potential extends widely across various fields of science and engineering that involve complex-valued data, such as wireless communications and quantum mechanics.
CommentsSubmitted to IEEE Trans. ASLP