arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DyRA:用于深度神经网络中高效矩阵乘法的动态残差逼近

DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs

Daewon Chae, Hyunwon Chung, Changwoo Lee, Hun-Seok Kim

arXiv 2610.02882首次发表:更新:

发表机构

University of Michigan; Qualcomm AI Research(密歇根大学; 高通人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DyRA通过动态逼近并校正结构化权重逼近产生的输出残差误差,在相同计算预算下更忠实逼近完整矩阵乘法,在视觉、语音和语言模型中一致提升精度-效率权衡,并在DINOv3上实现1.5倍GPU加速和3倍精度下降减少。

AI 中文摘要

大规模基础模型在多种任务上表现出强大的性能,但其规模使得推理成本高昂,这在很大程度上归因于密集矩阵乘法。先前的工作通过用高效的结构化形式(如低秩分解)替换密集权重矩阵来降低这一成本。然而,这些方法逼近的是权重而非决定推理精度的输出激活。因此,小的权重空间误差可能被输入激活放大,产生大的输出误差。在这项工作中,我们提出了DyRA,一种输入自适应方法,通过在推理过程中校正残差输出误差来改进结构化矩阵乘法逼近。我们表明,通过直接优化输出的低秩因子,可以更有效地逼近矩阵乘法。DyRA基于这一见解,动态逼近并校正由结构化权重逼近引入的输出误差。这将高效的结构化计算与输入相关的校正相结合,在相同计算预算下产生对完整矩阵乘法更忠实的逼近。在视觉、语音和语言模型中,DyRA始终优于单独的结构化权重逼近,提高了精度-效率权衡。值得注意的是,与仅权重的基线相比,DyRA在DINOv3上实现了1.5倍的端到端GPU加速,同时将精度下降减少了3倍以上。

英文摘要

Large-scale foundation models achieve strong performance across diverse tasks, but their size makes inference costly, largely due to dense matrix multiplications. Prior work reduces this cost by replacing dense weight matrices with efficient structured forms such as low-rank factorizations. However, these methods approximate weights rather than the output activations that determine inference accuracy. Consequently, small weight-space errors can be amplified by input activations, producing large output errors. In this work, we propose DyRA, an input-adaptive method that improves structured matrix multiplication approximation by correcting residual output errors during inference. We show that matrix multiplication can be approximated more effectively by directly optimizing low-rank factors of the output. DyRA builds on this insight by dynamically approximating and correcting the output error introduced by structured weight approximations. This combines efficient structured computation with input-dependent correction, yielding a more faithful approximation of full matrix multiplication under the same computational budget. Across vision, speech, and language models, DyRA consistently improves the accuracy-efficiency trade-off over structured weight approximations alone. Notably, DyRA achieves a 1.5$\times$ end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than 3$\times$ relative to weight-only baselines.

CommentsNeurIPS 2026. Code: https://github.com/daewon88/DyRA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑