arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25963cs.LG

GeoPair:保持几何特性的跨层分解用于免训练Transformer压缩

GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression

  • MWS AI(MWS人工智能公司)
  • ITMO University(圣彼得堡国立信息技术机械与光学大学)

机构由 AI 辅助整理,请以论文原文为准。

Baher Mohammad, Ammar Ali, Stamatios Lefkimmiatis

中文总结 AI 辅助

提出免训练的跨层权重配对与共享字典分解框架,保留层几何结构,实现高效压缩并优于现有方法。

中文摘要 AI 辅助

Transformer架构表现出跨层的冗余性,然而训练后压缩流程通常孤立地优化各层,或依赖忽视层特定激活几何结构的启发式分组策略。我们提出了一种原则性的、免训练的框架,该框架顺序优化跨层权重配对和共享字典分解。我们的方法不是强制相邻层的权重共享一个基或启发式地合并激活统计量,而是识别结构上兼容的投影,并学习一个共享表示,该表示能更好地保留每层独特的校准几何结构。与结构化稀疏性相结合,这在不牺牲功能保真度的情况下产生了高效的权重分解。在多种架构、规模和多模态中,我们的方法取得了最先进的结果,始终优于独立的结构化权重分解和替代的成对权重分解(这些方法基于启发式分组策略)。通过用收敛的、优化驱动的流程取代启发式工程策略,我们为跨不同模态的可扩展Transformer压缩建立了理论基础。

英文摘要

Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations. Rather than forcing weights of adjacent layers to share a basis or heuristically merging activation statistics, our approach identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry. Coupled with structured sparsity, this yields highly efficient weight decompositions without sacrificing functional fidelity. Across diverse architectures, scales, and modalities, our method achieves state-of-the-art results, consistently outperforming independent structured weight decompositions and alternative pairwise weight factorizations, which operate under heuristic grouping strategies. By replacing heuristic engineering strategies with a convergent, optimization-driven pipeline, we establish a theoretically grounded foundation for scalable, transformer compression across different modalities.

↑