arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CHASE:面向几何感知模型工程的信道对齐结构利用

CHASE: Channel-Aligned Structure Exploitation for Geometry-Aware Model Engineering

Wei Wang, Wei Jiang, Ziran Liu

arXiv 2610.09476首次发表:更新:

发表机构

Futurewei Technologies; Shanghai Institute for Mathematics and Interdisciplinary Sciences (SIMIS)(未来科技公司; 上海数学与交叉科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出CHASE框架,利用几何与谱对齐(GSA)结构,涵盖六个应用,包括三种新方法CAGA、SAKV和CAPS,分别改进注意力头共享、KV缓存压缩和结构化剪枝,实验验证其有效性。

AI 中文摘要

几何与谱对齐(GSA)通过谱集中度、物理信道对齐、支撑结构以及奇异基的变化来刻画训练后的网络。在本文中,我们提出CHASE(信道对齐结构利用),以在实际模型设计中使用这些结构。CHASE涵盖模型修改、重构和压缩中的六个应用。CORA、COEC和CORAM将GSA应用于参数高效微调、结构化剪枝补偿和模型合并。我们进一步开发了三种新方法。CAGA使用GSA来识别可以共享KV表示的多头注意力头,并通过几何对齐和低秩子空间提取来构建共享的键和值头。SAKV使用GSA来确定哪些相邻层可以共享低秩KV缓存表示,以及每个层组保留的秩。CAPS使用GSA谱结构对输出神经元进行分组,并为每组分别选择保留的输入信道。CORA、COEC和CORAM的结果确立了GSA在适配、剪枝补偿和模型合并方面的有效性。在CAGA上的实验表明,几何共享头构建显著改善了MHA到GQA的转换,而SAKV和CAPS在KV缓存压缩和结构化剪枝方面优于代表性基线。这些结果表明,GSA识别的结构可直接用于设计一系列模型操作的方法。

英文摘要

Geometric and Spectral Alignment (GSA) characterizes trained networks through spectral concentration, physical-channel alignment, support structure, and changes in singular bases. In this paper, we propose CHASE (Channel-Aligned Structure Exploitation) to use these structures in practical model design. CHASE covers six applications across model modification, reconfiguration, and compression. CORA, COEC, and CORAM apply GSA to parameter-efficient finetuning, structured-pruning compensation, and model merging. We further develop three new methods. CAGA uses GSA to identify multi-head attention heads that can share a KV representation and constructs the shared key and value heads through geometric alignment and low-rank subspace extraction. SAKV uses GSA to determine which adjacent layers can share a low-rank KV-cache representation and the retained rank for each layer group. CAPS uses GSA spectral structure to group output neurons and selects retained input channels separately for each group. Results from CORA, COEC, and CORAM establish the effectiveness of GSA for adaptation, pruning compensation, and model merging. Experiments on CAGA show that geometric shared-head construction substantially improves MHA-to-GQA conversion, and SAKV and CAPS improve over representative baselines for KV-cache compression and structured pruning. These results show that the structures identified by GSA can be used directly to design methods for a range of model operations.

Comments26 pages, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑