减少矩阵乘法:面向大语言模型推理的输入自适应矩阵乘积约简方法
Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference
查看机构详情
- University of Chicago(芝加哥大学)
- Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究针对Transformer语言模型推理开销高的问题,提出无需训练的输入自适应RMM方法,通过选择矩阵乘积信息切片减少计算,在多任务多模型上验证其鲁棒性,可提升长序列推理效率,且适用于多模态场景。
中文摘要 AI 辅助
基于Transformer的语言模型性能优异,但因反复进行高维矩阵乘法产生了可观的推理开销。我们提出Reduced Matrix Multiplication(RMM,减少矩阵乘法),一种无需训练、输入自适应的推理方法,通过沿Transformer矩阵乘积的收缩维度选择信息切片来减少计算量,且无需修改模型权重。在简单的保留率控制下,RMM可实现平滑且可预测的准确率-效率权衡。在参数规模从10亿到700亿的各类语言模型上,我们发现约简容忍度取决于模型家族、任务、组件及保留率,且通常随模型规模增大而提升。在适度约简下,RMM在评估的判别式任务、自回归生成任务及长上下文设置中均保持鲁棒性。我们进一步证明该原理可扩展至多模态视觉-语言推理。机制消融实验揭示了Transformer内部的结构不对称性:注意力侧计算的可约性远高于MLP组件。最后,在NVIDIA A100上使用自定义内核进行的实际运行时间基准测试表明,这些计算节省可转化为实际的运行时增益,尤其在长序列场景下效果显著。综上,这些结果表明RMM是输入自适应推理时优化的可扩展方向。
英文摘要
Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. We propose Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without modifying model weights. Under a simple retention-ratio control, RMM provides a smooth and predictable accuracy-efficiency trade-off. Across language models ranging from 1B to 70B parameters, we find that reduction tolerance depends on the model family, task, component, and retention ratio, although it often improves with model scale. Under moderate reduction, RMM remains robust across the evaluated discriminative, autoregressive generation, and long-context settings. We further show that the same principle extends to multimodal vision-language inference. Mechanistic ablations reveal a structural asymmetry within Transformers: attention-side computations are substantially more reducible than MLP components. Finally, wall-clock benchmarks with custom kernels on an NVIDIA A100 show that these computational savings can translate into practical runtime gains, especially at longer sequence lengths. Together, these results position RMM as a scalable direction for input-adaptive inference-time optimization.