发表机构
Columbia University; Recognition Technologies, Inc.(哥伦比亚大学; 识别技术公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对边缘设备音高估计的约束,将频率卷积网络分解为低秩形式,参数减少35.9%且精度不降,并引入噪声范围增强,同时开发了高效C内核,显著提升推理速度。
AI 中文摘要
边缘设备上的音高估计同时受到三个方面的限制:模型必须小,在输入有噪声时必须保持准确,并且一帧必须在帧周期内产生。在本报告中,我们早期工作中的频率卷积网络(FrCN)被分解为低秩形式。参数数量减少了35.9%,从17,787个降至11,397个,并且在领域内或两个模型从未训练过的语料库上,准确率均未下降。训练期间使用的噪声范围也被证明在设置模型在严重干扰下的行为方面占主导地位。当训练噪声下限从+6.02 dB降至-20 dB时,干净RPA50损失了0.35个百分点,而-20 dB时的RPA50从2.44提高到22.41。这一效应比所测量的任何架构效应大约大两个数量级。TDNN-F中使用的半正交约束被发现与瓶颈内的归一化冗余,当两者同时应用时准确率会下降。分解是否节省时间取决于运行时:在急切PyTorch中,分解后的模型慢45%,而在编译内核中则快15%。为了部署,编写了一个小型C内核。它需要83 kB的磁盘空间,并且除了libc和libm之外不需要任何运行时库。在所有测试的四个CPU上,它都比OpenBLAS快,比ONNX Runtime快3.8倍,比PyTorch快13倍。其输出在每模型271,893个保留帧上与PyTorch进行了核对,并且在每一帧上都选择了相同的音高区间。
英文摘要
Pitch estimation on an edge device is constrained in three ways at once. The model must be small, it must stay accurate when the input is noisy, and one frame must be produced inside the frame period. In this report the Frequency Convolution Network (FrCN) of our earlier work is factored into a low-rank form. The number of parameters is reduced by 35.9%, from 17,787 to 11,397, and accuracy is not reduced, either in domain or on two corpora the model was never trained on. The range of the noise used during training is also shown to dominate the architecture in setting how the model behaves when the interference is severe. When the training noise floor is lowered from +6.02 dB to -20 dB, 0.35 points of clean RPA50 are lost and RPA50 at -20 dB is raised from 2.44 to 22.41. This effect is about two orders of magnitude larger than any architectural effect that was measured. The semi-orthogonal constraint used in TDNN-F is found to be redundant with the normalization inside the bottleneck, and accuracy is reduced when both are applied. Whether the factorization saves time depends on the runtime: in eager PyTorch the factored model is 45% slower, while in a compiled kernel it is 15% faster. For deployment, a small C kernel was written. It needs 83 kB on disk and no runtime library beyond libc and libm. It is faster than OpenBLAS on all four CPUs that were tested, faster than ONNX Runtime by 3.8 times, and faster than PyTorch by 13 times. Its output was checked against PyTorch on 271,893 held-out frames per model, and the same pitch bin was selected on every one of them.
Comments11 pages, 11 tables. Columbia University Nonlinear Control Laboratory Technical Report CUNLC-20260919-02. A shortened version has been submitted to ICASSP 2027