LoopCD:用于提升循环语言模型推理能力的循环级对比解码
LoopCD: Loop-wise Contrastive Decoding for Improving Reasoning in Looped Language Models
- Pohang University of Science and Technology (POSTECH)(浦项科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对循环语言模型推理中的循环不稳定性问题,提出循环级对比解码LoopCD,通过对比早期与最终迭代logits干预困难标记,无需训练且开销极小,有效提升多种推理任务性能。
AI中文摘要:
循环语言模型(LoopLMs)通过使用共享权重递归地细化内部潜在表征来执行“潜在推理”,为显式言语推理提供了一种更有效的替代方案。尽管其有效,我们发现循环语言模型仍易受循环不稳定性影响:跨迭代的不稳定细化可能产生与推理错误相关的局部不确定“困难”标记。为解决此问题,我们提出LoopCD,一种循环级对比解码方法,通过在推理时干预这些标记来增强循环语言模型的推理性能。具体而言,我们利用循环语言模型的内部动态,将早期迭代的logits与最后一次细化迭代的logits进行对比,以形成最终采样分布。我们发现该策略非常高效,仅引入可忽略的推理开销且无需额外训练,同时通过自然细化推理关键困难标记有效提升推理性能。大量实验表明,我们的方法在各种推理任务上提升了近期代表性循环语言模型的性能。
英文摘要:
Looped Language Models (LoopLMs) perform "latent reasoning" by recursively refining internal latent representations with shared weights, offering a more effective alternative to explicit verbal reasoning. Despite their effectiveness, we find that LoopLMs remain prone to loop instability: unstable refinement across iterations can produce localized uncertain "hard" tokens associated with reasoning errors. To address this, we propose LoopCD, loop-wise contrastive decoding that enhances the reasoning performance of LoopLMs by intervening on these tokens at inference time. Specifically, we exploit the internal dynamics of LoopLMs and contrast the logits from earlier iterations with logits from the last refined iteration to form the final sampling distribution. We find that this strategy is highly efficient, introducing only negligible inference overhead and requiring no additional training, while effectively improving reasoning performance by naturally refining reasoning-critical hard tokens. Extensive experiments show that our method improves the performance of recent representative LoopLMs across various reasoning tasks.