AI 中文总结
本文提出AudioTQ,一种基于随机哈达玛旋转与劳埃德-麦克斯韦量化的6位CPU时域音频编解码器,无需硬件加速即可实现实时执行,可将音频文件大小缩减74.4%,SQNR约30 dB。
AI 中文摘要
传统有损音频压缩算法依赖心理声学模型与频域表示(如MP3、AAC和Opus)来丢弃人类听觉系统无法感知的信息。这些方法虽高效,但计算复杂且具有领域特异性。本文提出AudioTQ的设计与数学公式,这是一种直接在时域运行的无数据感知有损音频编解码器。受大语言模型(LLM)权重量化技术(特别是TurboQuant框架)启发,AudioTQ采用正交随机快速沃尔什-哈达玛变换(FWHT)旋转,将波动的时域幅度统一为可预测的标准正态分布。这使得可使用离线训练的均方误差(MSE)最优6位劳埃德-麦克斯韦量化器进行坐标级标量量化,并辅以1比特量化联合最小二乘(QJL)残差校正层。生成的7位虚拟索引被打包为原生8位容器,与标准CPU寄存器边界对齐,以确保无需硬件并行加速器即可实现实时单线程执行。本文详述24位录音室音源的逐位重构,分析FWHT的蝶形网络,推导稀疏输入下的数学失效模式,并呈现基准测试结果:物理大小缩减高达74.4%,同时量化信噪比(SQNR)约为30 dB。
英文摘要
Lossy audio compression algorithms traditionally rely on psychoacoustic modeling and frequency-domain representations (e.g., MP3, AAC, and Opus) to discard information that is imperceptible to the human auditory system. While highly effective, these approaches are computationally complex and domain-specific. In this paper, we present the design and mathematical formulation of AudioTQ, a data-oblivious lossy audio codec that operates directly in the time domain. Inspired by Large Language Model (LLM) weight quantization techniques (specifically the TurboQuant framework), AudioTQ uniformizes volatile time-domain amplitudes into a predictable standard normal distribution using an orthonormal, randomized Fast Walsh-Hadamard Transform (FWHT) rotation. This enables coordinate-wise scalar quantization using an offline-trained, MSE-optimal 6-bit Lloyd-Max quantizer, augmented by a 1-bit Quantized Joint Least-Squares (QJL) residual correction layer. The resulting 7-bit virtual indices are packed into native 8-bit containers, aligning with standard CPU register boundaries to ensure real-time single-threaded execution without hardware parallel accelerators. We detail the bitwise reconstruction of 24-bit studio stems, analyze the butterfly network of the FWHT, derive the mathematical failure modes under sparse inputs, and present benchmarks showing up to 74.4% physical size reduction alongside a Signal-to-Quantization-Noise Ratio (SQNR) of ~30 dB.
Comments8 pages, 1 figure, 1 table. Code is available at https://github.com/lostmartian/audioTQ