arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

因果门控:用于Transformer模块剪枝的因果重要性蒸馏

CausalGate: Causal Importance Distillation for Transformer Module Pruning

Kiran Nair, Smriti Regmi, Rodrigue Rizk

arXiv 2607.22720首次发表:更新:

AI 中文总结

研究针对大语言模型自适应推理方法不足,提出CausalGate框架,在校准阶段隔离子层测量语义损伤,再用特定目标和损失将结构重要性提炼为标量门控,经实验验证其在多模型上优于基线,能降低硬件延迟且无操作开销。

AI 中文摘要

现有的大语言模型自适应推理方法依靠观测启发式方法(如隐藏状态相似度或激活幅度)来丢弃冗余模块。然而,这些基于相关性的指标往往无法捕捉对语义准确性至关重要的微妙非线性结构计算。我们引入了因果门控(CausalGate),这是一个用于高效计算的Transformer推理的干预引导框架。在校准阶段,因果门控隔离各个注意力和MLP子层,将其输出归零,并通过最终对数分布的库尔贝克-莱布勒散度测量确切的语义损伤。为消除运行时路由开销,利用指数移动平均平滑目标和可微成对排序损失,将这种结构重要性层次提炼为一组全局静态轻量级标量门控。在跨语言建模和常识推理基准测试的TinyLlama-1.1B、Qwen2.5-3B和Llama-3.1-8B上进行评估,因果门控始终优于突出的动态路由和层跳过基线,将理论计算节省转化为具体硬件延迟降低,且无操作开销。

英文摘要

Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop redundant modules. However, these correlation-based metrics often fail to capture subtle, non-linear structural computations vital for semantic accuracy. We introduce CausalGate, an intervention-guided framework for compute-efficient transformer inference. During a calibration phase, CausalGate isolates individual Attention and MLP sub-layers, zeros out their respective outputs, and measures the exact semantic damage via the Kullback-Leibler divergence of the final logit distribution. To eliminate runtime routing overhead, this structural importance hierarchy is distilled into a global set of static, lightweight scalar gates using an Exponential Moving Average smoothing objective paired with a differentiable pairwise ranking loss. Evaluated on TinyLlama-1.1B, Qwen2.5-3B, and Llama-3.1-8B across language modeling and commonsense reasoning benchmarks, CausalGate consistently outperforms prominent dynamic routing and layer-skipping baselines, translating theoretical compute savings into concrete hardware latency reductions with zero operational overhead.

CommentsIn-Review at a conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑