发表机构
POSTECH; Daegu Gyeongbuk Institute of Science and Technology (DGIST); Hanyang University(浦项科技大学; 大邱庆北科学技术院; 汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对剪枝后大语言模型的文本退化问题,提出FOCUS与RePAIR两种令牌级引导微调方法,在两类生成任务上均有效减少重复并提升生成质量。
AI 中文摘要
剪枝是压缩大语言模型(LLMs)的实用方法,但即便困惑度和任务准确率基本保持不变,它仍会加剧文本退化,尤其是重复循环现象。本研究将解码过程视为进入并持续处于一小部分循环上下文的动力学过程,对该失效模式进行令牌级分析,将退化分解为循环进入风险与循环持续性,且表明持续性由令牌采样集中分配给合理替代项的逃逸质量控制。基于这些发现,我们提出两项用于剪枝后微调的令牌级引导目标:FOCUS将蒸馏权重重新分配至高置信度教师区域以抑制泄漏,RePAIR利用以起始为中心的正负延续对及边际损失,以推广合理替代项并防止过早陷入重复循环。在开放式延续与指令生成任务上的实验表明,两种方法均能持续减少重复并提升生成质量。
英文摘要
Pruning is a practical approach to compress large language models (LLMs), but it can amplify text degeneration, especially repetition loops, even when perplexity and task accuracy remain largely unchanged. In this work, we present a token-level analysis of this failure mode by viewing decoding as a dynamical process that enters and persists in a small set of recurrent contexts. Our analysis decomposes degeneration into loop entry risk and loop persistence, and shows that persistence is controlled by the escape mass assigned to plausible alternatives within the token sampling set. Motivated by these findings, we propose two token-level guidance objectives for post-pruning fine-tuning. FOCUS reweights distillation toward high-confidence teacher regions to suppress leakage, while RePAIR uses onset-centered positive/negative continuation pairs with a margin loss to promote plausible alternatives and prevent early commitment to repetition loops. Experiments on open-ended continuation and instruction-based generation show that both methods consistently reduce repetition and improve generation quality.
CommentsAccepted to ICML 2026 as a Spotlight