RATIO:面向量化推理模型的推理分析与令牌级推理优化
RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models
浏览论文内容
中文总结 AI 辅助
提出RATIO框架,通过量化感知行为分析识别过度思考令牌并施加定制惩罚,在量化推理模型上提升精度达9.8个百分点,缩短思维链长度达51.3%。
中文摘要 AI 辅助
后训练量化(PTQ)已成为降低大型语言模型(LLMs)内存占用和推理成本的广泛采用的技术。然而,近期研究表明,当应用于推理模型时,PTQ不仅降低推理性能,还加剧过度思考,导致更长的推理轨迹。这些问题可能抵消低精度推理预期的效率提升。现有方法主要依赖复杂的优化过程。更近期的轻量级推理策略则使用预定义的过度思考标记,限制了它们在量化模型中的适应性。为解决这些问题,我们提出了推理分析与令牌级推理优化(RATIO),一个识别模型特定过度思考令牌并为其分配定制惩罚的框架。RATIO首先引入量化感知推理行为分析(QRBA),通过分析全精度与量化模型之间的差异来识别过度思考令牌。然后采用令牌特定惩罚确定(TSPD),利用全精度指导推导令牌特定惩罚,无需额外训练。大量实验表明,RATIO相比现有令牌级干预实现了更好的精度-效率权衡。具体而言,与量化基线相比,RATIO实现了高达9.8个百分点的精度提升,并将思维链(CoT)长度减少了高达51.3%。代码将在该https URL提供。
英文摘要
Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). However, recent studies reveal that when applied to reasoning models, PTQ not only degrades reasoning performance but also exacerbates overthinking, leading to longer reasoning trajectories. These issues may offset the efficiency gains expected from lower-precision inference. Existing approaches mainly rely on complex optimization procedures. More recent lightweight inference strategies instead use predefined overthinking markers, limiting their adaptability across quantized models. To address these issues, we propose Reasoning Analysis and Token-level Inference Optimization (RATIO), a framework that identifies model-specific overthinking tokens and assigns each a tailored penalty. RATIO first introduces Quantization-aware Reasoning Behavior Analysis (QRBA) to identify overthinking tokens by analyzing discrepancies between full-precision and quantized models. It then adopts Token-Specific Penalty Determination (TSPD), which leverages full-precision guidance to derive token-specific penalties without additional training. Extensive experiments show that RATIO achieves a better accuracy-efficiency trade-off than existing token-level interventions. Specifically, RATIO achieves up to 9.8 points accuracy improvement and reduces chain-of-thought (CoT) length by up to 51.3% compared with quantized baselines. The code will be available at https://github.com/steven-bao1/RATIO.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。