arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18723cs.AIcs.LG

超越截断:将LLM解码重新思考为集成剪枝

Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning

Dunyao Xue, Chengshuo Du, Zhengbo Wang, Wenlin Dai, Cheng Meng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出ME-Decoding,将LLM解码视为集成剪枝,用马氏距离目标优化子集选择,通过自适应核构建相似度矩阵折扣冗余路径,并设计近线性贪心算法,实现低开销的稳健解码。

中文摘要 AI 辅助

我们提出了马氏距离集成解码(ME-Decoding),一种新颖的大语言模型(LLM)解码框架,它将候选词元选择视为集成剪枝问题。现有的选择策略主要依赖标量概率,忽略了几何语义关系,导致候选冗余。同时,当前的几何感知方法通常需要复杂的优化或直接重新加权原始词元概率,造成显著的计算开销或推理不稳定性。为解决这一问题,我们将解码形式化为一个子集优化问题,使用基于马氏距离的目标函数来增强语义多样性,同时保持高概率。具体而言,我们通过一个基于词元嵌入的自适应带宽核构建的词元相似度矩阵,动态地对冗余生成路径进行折扣。我们进一步设计了一种高效的贪心选择算法,在提前停止条件下,其复杂度在候选规模上接近线性,并建立了其理论近似保证。这使得ME-Decoding成为一个稳健、即插即用的模块,推理开销可忽略不计。在多种推理和生成任务上的广泛实验表明,我们的方法持续取得了强劲的性能。

英文摘要

We introduce Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding framework that frames candidate token selection as ensemble pruning. Existing selection strategies rely predominantly on scalar probabilities, ignoring geometric semantic relationships and causing candidate redundancy. Meanwhile, current geometry-aware methods often require complex optimization or directly reweighting the original token probabilities, leading to significant computational overhead or inference instability. To address this, we formulate decoding as a subset optimization problem using a Mahalanobis distance-driven objective to enhance semantic diversity while preserving high probabilities. Specifically, we dynamically discount redundant generation paths using a token similarity matrix, constructed via an adaptive-bandwidth kernel over token embeddings. We further devise an efficient greedy selection algorithm with near-linear complexity in the candidate size under early stopping, while establishing its theoretical approximation guarantees. This renders ME-Decoding a robust, plug-and-play module with negligible inference overhead. Extensive experiments across diverse reasoning and generation tasks demonstrate that our method consistently achieves strong performance.

发表机构

  • Institute of Statistics and Big Data, Renmin University of China(中国人民大学统计与大数据研究院)
  • Big Data and Responsible Artificial Intelligence for National Governance, Renmin University of China(中国人民大学大数据与国家治理研究院)
  • Center for Applied Statistics, Institute of Statistics and Big Data, Renmin University of China(中国人民大学统计与大数据研究院应用统计中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑