发表机构
Aalto University; Nokia; Nanyang Technological University(阿尔托大学; 诺基亚; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MU-MIMO神经接收机,研究4位浮点微格式(FP4)量化与剪枝,相比INT4显著减少性能损失,保持优于经典接收机,并大幅降低计算和存储成本。
AI 中文摘要
神经接收机性能优于传统5G NR处理链,但其计算和内存需求阻碍了实时部署。对于符合标准的MU-MIMO神经接收机,4位数字格式(而非仅位宽)决定了压缩是否能保持对经典接收机的增益。我们分别应用权重和激活量化感知训练(QAT)以及50%幅度剪枝,比较INT8/INT4和FP8(E4M3)/FP4(E2M1)权重与INT8后ReLU激活。在3GPP UMi信道上训练,并在TDL-B和TDL-C上评估,8位权重-激活模型在10%和1%误块率(BLER)下与FP32相差在0.05 dB以内。在4位时,均匀INT4损失3.3-3.7 dB并低于LS-LMMSE,而FP4将损失减半以上(1.3-1.4 dB),即使在剪枝后仍优于LS-LMMSE约0.5 dB。FP4更密集的接近零网格与训练权重分布匹配,且FP4避免了INT4中出现的残差路径过度剪枝。分析成本模型预测剪枝后的4位权重推理可减少66倍位操作和8.8倍权重存储。
英文摘要
Neural receivers outperform conventional 5G NR processing chains, but their compute and memory demands hinder real-time deployment. For a standard-compliant multi-user MIMO neural receiver, the 4-bit number format, not merely the bit width, determines whether compression preserves the gain over classical receivers. We apply weight and activation quantization-aware training (QAT) and, separately, 50% magnitude pruning, comparing INT8/INT4 and FP8 (E4M3)/FP4 (E2M1) weights with INT8 post-ReLU activations. Trained on 3GPP UMi channels and evaluated on TDL-B and TDL-C, 8-bit weight-activation models remain within 0.05 dB of FP32 at 10% and 1% block error rate (BLER). At 4 bits, uniform INT4 loses 3.3-3.7 dB and falls below LS-LMMSE, whereas FP4 more than halves this loss (1.3-1.4 dB) and still outperforms it by about 0.5 dB, even after pruning. FP4's denser near-zero grid matches the trained weight distribution, and FP4 avoids the residual-path over-pruning seen with INT4. An analytic cost model projects 66x fewer bit-operations and 8.8x less weight storage for pruned 4-bit-weight inference.
CommentsThis work has been submitted to the 2027 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2027)