arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LowRankArena:面向基于SVD的大语言模型压缩的标准化评估平台

LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

Zishan Shao, Lixun Zhang, Kangning Cui, Wenhao Wu, Jinhee Kim, Yixiao Wang, Ting Jiang, Hancheng Ye, Qinsi Wang, Fan Yang, Danyang Zhuo, Yiran Chen, Hai Li

arXiv 2608.26389首次发表:更新:

发表机构

Duke University; Wake Forest University(杜克大学; 维克森林大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出LowRankArena这一标准化评估平台,解决基于SVD的LLM压缩研究中评估不统一的问题,通过审计五种SVD方法揭示过往发现的条件性。

AI 中文摘要

基于奇异值分解(SVD)的低秩压缩已成为降低大语言模型(LLM)内存与计算成本的快速发展方向。然而,现有研究间的有意义对比仍存在困难,因为过往评估使用了不同的基准、不一致的压缩率及多样的设置,常无法将低秩效应与辅助技术分离。因此,目前仍不清楚报告的提升是否反映了方法层面的改进,还是评估协议的差异。这种可比性的缺失凸显了统一、可复现的评估平台的必要性。为解决该问题,我们提出LowRankArena,一个面向基于SVD的LLM压缩的标准化评估平台。LowRankArena统一了任务版本、均匀精度压缩预算、对比机制及推理测量,还提供了可复现的流水线,附带超过3 TiB的已发布压缩检查点。利用LowRankArena,我们对五种代表性SVD方法的对齐审计显示,过往发现在标准化协议下具有高度条件性:清晰的领先者与性能层级会随骨干模型与压缩率变化;多选择准确率可能掩盖困惑度的大幅下降;名义上的低秩节省会产生依赖工作负载且通常有限的端到端加速。我们的代码可在以下网址获取:this https URL。

英文摘要

SVD-based low-rank compression has become a fast-growing direction for reducing the memory and computational cost of large language models (LLMs). However, meaningful comparison across existing studies remains difficult as prior evaluations use varied benchmarks, inconsistent ratios, and diverse setups, often failing to isolate low-rank effects from auxiliary techniques. As a result, it remains unclear whether reported gains reflect method-level improvements or differences in evaluation protocol. This lack of comparability highlights the need for a unified, reproducible evaluation platform. To address this problem, we present LowRankArena, a standardized evaluation platform for SVD-based LLM compression. LowRankArena unifies task versions, uniform-precision compression budgets, comparison regimes, and inference measurements, and provides a reproducible pipeline with over 3 TiB released compressed checkpoints. Using LowRankArena, our aligned audit of five representative SVD methods reveals that prior findings are highly conditional under standardized protocols: clear leaders and performance tiers shift across backbones and keep ratios, multiple-choice accuracy can hide large perplexity degradation, and nominal low-rank savings yield workload-dependent and often limited end-to-end speedups. Our code is available at: https://github.com/Zishan-Shao/lowrankarena.git.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑