发表机构
Chalmers University of Technology(查尔姆斯理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出D-SLR分解,作为截断SVD的闭式替代方法,通过将行限制为逐字存储或低秩近似,在相同或更少参数下提升重构性能,无需调参,在多类真实数据实验中验证了其有效性。
AI 中文摘要
用于重构的矩阵压缩目前仍默认采用截断奇异值分解(truncated SVD),以单一低秩结构近似数据。常见做法是通过添加重叠行稀疏分量进一步降低残差,但求解该联合问题的方法通常需要迭代求解器并调整正则化参数。我们提出不相交行稀疏加低秩(Disjoint Row-Sparse plus Low-Rank,D-SLR)分解,这是一种可直接替代截断奇异值分解的闭式方法,其性能优于或完全匹配截断奇异值分解。D-SLR将行限制为要么逐字存储,要么由低秩拟合近似,二者不可兼得。在平方误差准则下,该限制无性能损失:在任意非平凡秩和存储行数(形状)下,均可通过更少参数实现联合最优解。当存储行数为零时,D-SLR退化为截断奇异值分解,因此在相同成本下其性能绝不会更差。该算法会遍历整个误差-参数权衡空间,后续可通过提供的误差目标、参数数量或选择规则来确定解。整个网格及求解过程仅需三次奇异值分解,无需任何调参或正则化。我们推导了无假设的后验误差下界,适用于任意形状,为每个解提供了可计算的证书,用于评估选择其他秩和存储行数的潜在收益。在合成数据及真实数据(大语言模型嵌入表、网络流量、高光谱图像)上的实验证实了该方法的收益,并对证书进行了量化。
英文摘要
Compressing a matrix for reconstruction still defaults to the truncated SVD, approximating the data with a single low-rank structure. It is common to reduce the residual further by adding an overlapping row-sparse component, but methods that solve this joint problem often require iterative solvers and tuning of regularization parameters. We propose the Disjoint Row-Sparse plus Low-Rank (D-SLR) decomposition, a closed-form drop-in for the truncated SVD that improves or exactly matches it. D-SLR restricts rows to either being stored verbatim or approximated by the low-rank fit, never both. Under squared error this restriction costs nothing: the joint optimum is attainable disjointly with fewer parameters at every non-trivial rank and stored row count (shape). With zero stored rows D-SLR reduces to the truncated SVD, so it never does worse at equal cost. The algorithm scores the entire error-versus-parameters tradeoff, and the solution is chosen afterwards by a supplied error target or parameter count, or by a selection rule. The grid and solution together cost three SVDs, with no tuning or regularization. We derive an assumption-free, a-posteriori lower bound on the error at every shape, giving each solution a computable certificate on the potential gain of any other choice of rank and stored rows. Experiments on synthetic and real data (LLM embedding tables, network traffic, hyperspectral images) confirm the gains and quantify the certificate.