发表机构
Santa Clara University; Sogang University(圣塔克拉拉大学; 西江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型中如何在保留模型质量时减少计算,提出灵敏度感知阈值处理方法SATS及轻量级令牌路由框架,经实验评估,SATS在匹配稀疏度时优于基线,令牌路由有更好的质量-吞吐量权衡,改进方法可提升大语言模型权衡效果。
AI 中文摘要
大语言模型中的高效推理需要在保留模型质量的同时决定何处可以减少计算。我们通过多层感知器(MLP)激活稀疏化和令牌级条件路由来研究这个问题。首先提出用于稀疏性的灵敏度感知阈值处理(SATS),一种使用局部MLP输出灵敏度代理选择逐层门限阈值的阈值校准方法,而非直接从激活百分位数校准阈值。还引入轻量级令牌路由框架,在每个令牌基础上动态选择基本路径和修改路径。在多个开放权重的大语言模型上评估这两种方法,结果表明SATS在匹配实际稀疏度时优于基于阈值的稀疏化基线,令牌路由比静态激活修改基线产生更有利的质量-吞吐量权衡。总体而言,改进的阈值校准和令牌路由可改善大语言模型中的质量-吞吐量权衡。
英文摘要
Efficient inference in Large Language Models (LLMs) requires deciding where computation can be reduced while preserving model quality. We study this problem through multilayer perceptron (MLP) activation sparsification and token-level conditional routing. We first propose Sensitivity-Aware Thresholding for Sparsity (SATS), a threshold calibration method to choose layerwise gate thresholds using a local MLP output sensitivity proxy rather than calibrating thresholds directly from activation percentiles. While SATS retains the existing mechanism of sparsifying MLP activations by thresholding gate activations, it replaces percentile-based calibration with a sensitivity-aware selection rule. We then introduce a lightweight token routing framework that dynamically selects between a base path and a modified path on a per-token basis, rather than applying the modified computation uniformly to all tokens. We evaluate both methods on multiple recent open-weight LLMs. Our results show that SATS improves over the threshold-based sparsification baseline at matched actual sparsity and that token routing yields a more favorable quality-throughput trade-off than static activation modification baselines. Overall, our results suggest that improved threshold calibration and token routing can improve the quality-throughput trade-off in LLMs.