推理感知压缩:识别并保护脆弱推理电路以实现节能的大语言模型部署
Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment
- Carnegie Mellon University Africa(卡内基梅隆大学非洲校区)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出推理感知压缩框架,通过识别并保护脆弱推理电路,在五个推理基准上实现比统一量化更优的能耗与精度帕累托最优。
AI中文摘要:
大型推理模型(LRMs)在部署过程中产生大量能源成本,然而当前的压缩方法对所有组件采用统一量化,可能损害关键的推理电路。我们提出了一种推理感知的压缩框架,该框架在五个推理基准(GSM8K、FOLIO、MATH-500、ProofWriter和MuSiQue)上对量化条件进行基准测试,并进行硬件级GPU能量测量;通过在一个保留的校准分割上进行扰动扫描,对全部196-224个(层,投影)对进行每模块INT4脆弱性分析,然后选择性地将最敏感的电路恢复为FP16。我们发现了三个结果。首先,INT4量化可能通过延长推理链来增加能量消耗;在GSM8K上,25%的功率降低变成了净能量增加。其次,脆弱性与任务相关:注意力投影对数学推理更为关键,而在逻辑推理中,敏感性模式因架构而异。第三,选择性压缩实现了统一方法无法达到的帕累托最优点:R1-Qwen-7B在ProofWriter上的前10%在-9.7%能量下比FP16提高了+12个百分点,并在五个推理基准的保留数据上得到验证。
英文摘要:
Large Reasoning Models (LRMs) impose substantial energy costs during deployment, yet current compression methods apply uniform quantization across all components, risking damage to critical reasoning circuits. We present a reasoning-aware compression framework that benchmarks quantization conditions across five reasoning benchmarks, GSM8K, FOLIO, MATH-500, ProofWriter, and MuSiQue, with hardware-level GPU energy measurement; profiles per-module INT4 vulnerability across all 196-224 (layer, projection) pairs via a perturbation sweep on a held-out calibration split, then selectively restores the most sensitive circuits to FP16. Three findings emerge. First, INT4 quantization can increase energy by extending reasoning chains; a 25% power reduction becomes a net energy increase on GSM8K. Second, vulnerability is task-dependent: attention projections are more critical for mathematical reasoning, and sensitivity patterns differ by architecture in logical inference. Third, selective compression achieves Pareto-optimal points inaccessible to uniform methods: R1-Qwen-7B Top-10% on ProofWriter gains +12 pp over FP16 at -9.7% energy, validated on held-out data across five reasoning benchmarks.