Floating-Point Microformat Quantization and Pruning for Efficient MU-MIMO Neural Receivers
En palabras de los autores
Neural receivers outperform conventional 5G NR processing chains, but their compute and memory demands hinder real-time deployment. For a standard-compliant multi-user MIMO neural receiver, the 4-bit number format, not merely the bit width, determines whether compression preserves the gain over classical receivers. We apply weight and activation quantization-aware training (QAT) and, separately, 50% magnitude pruning, comparing INT8/INT4 and FP8 (E4M3)/FP4 (E2M1) weights with INT8 post-ReLU activations. Trained on 3GPP UMi channels and evaluated on TDL-B and TDL-C, 8-bit weight-activation models remain within 0.05 dB of FP32 at 10% and 1% block error rate (BLER). At 4 bits, uniform INT4 loses 3.3-3.7 dB and falls below LS-LMMSE, whereas FP4 more than halves this loss (1.3-1.4 dB) and still outperforms it by about 0.5 dB, even after pruning. FP4's denser near-zero grid matches the trained weight distribution, and FP4 avoids the residual-path over-pruning seen with INT4. An analytic cost model projects 66x fewer bit-operations and 8.8x less weight storage for pruned 4-bit-weight inference.
Apareció: lunes, 28 de septiembre. arXiv. Preprint, todavía sin revisión por pares.
Comentario de los autores: This work has been submitted to the 2027 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2027)