arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02083cs.LGcs.CL

XMerge:用于大语言模型深度压缩的跨轴选择与重构层合并方法

XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression

Jundong Hu, Shekar Ramachandran

首次发表
浏览论文内容

中文总结 AI 辅助

XMerge是一种无需任务标签或端到端微调的LLM深度压缩后处理方法,通过跨轴选择和局部边界重构实现高效压缩,在7种主干模型及多个基准上表现优异,仅需数万个请求即可收回构建成本。

中文摘要 AI 辅助

移除完整的Transformer层可保留标准的服务架构,但现有的深度压缩方法会导致大量质量损失,且该损失在不同模型间存在不可预测的差异。我们提出XMerge,这是一种仅需训练后处理的方法,包含两个组件:跨轴选择用于识别具有低相对幅度和低隐状态角度变化的块,局部边界重构则重新拟合相邻的保留块,使其匹配原始两个块的输出。XMerge无需任务标签或端到端微调,既不引入架构变更,也不增加推理时的参数。在7种Llama和Qwen主干模型(0.5B至8B)、5种已发表的基准方法及3种层缩减级别下,其在最激进的层移除场景中对基准方法的优势最大:当k=4时,在CORE(22个任务的 aggregate 基准)上,7种主干模型中有6种排名第一;在MMLU基准上,7种主干模型中有6种排名第一(同时在两个基准上排名第一的有5种),同时避免了多个竞争算子出现的困惑度大幅上升问题。在任务级自助法检验中,CORE的三个最大优势的95%置信区间不包含零,其余优势则与平局一致。在14个(模型、缩减级别)组合中,它也是所有被评估算子中唯一未出现崩溃的,在零样本和上下文学习场景中均排名前2;在首个校准探针(一种主干模型)上,它是校准效果最佳的算子。 ablation 实验显示,局部重构提供了大部分性能提升,而跨轴融合在两个选择轴意见不一致时发挥作用。额外的构建成本可在大约数万个请求后通过每token解码的节省收回。

英文摘要

Removing complete transformer layers preserves a standard serving architecture, but existing depth-compression methods can lose substantial quality, and the loss varies unpredictably across models. We introduce XMerge, a post-training method with two components. Cross-axis selection identifies a block with low relative-magnitude and angular hidden-state change, and local boundary reconstruction re-fits the adjacent surviving block to match the original two-block output. XMerge uses no task labels or end-to-end fine-tuning, and it introduces neither architectural changes nor additional inference-time parameters. Across seven Llama and Qwen backbones (0.5B-8B), five published baselines, and three layer-reduction levels, its advantage over baselines is largest at the most aggressive removal: at k=4 it ranks first on six of seven backbones on CORE (a 22-task aggregate) and, separately, on six of seven on MMLU (five of seven on both at once), while avoiding the large perplexity increases of several competing operators. In a task-level bootstrap, the 95% confidence intervals for the three largest CORE margins exclude zero; the remaining margins are consistent with ties. Across the 14 (model, regime) cells it is also the only evaluated operator that never collapses, ranking top-2 in both zero-shot and in-context regimes; on a first calibration probe (one backbone) it is the best-calibrated operator. Ablations show that local reconstruction provides most of the gain, while cross-axis fusion helps when the two selection axes disagree. The additional construction cost is recovered through per-token decode savings after roughly tens of thousands of requests.

发表机构

  • PayPal AI(PayPal人工智能部门)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑