发表机构
The Ohio State University; Princeton University(俄亥俄州立大学; 普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CritICL是一种新型推理时框架,利用小语言模型的结构化失败模式,通过两种变体提升推理性能,效率优于标准上下文学习和测试时缩放方法。
AI 中文摘要
近期推理时缩放的进展显著提升了大语言模型(LLM)的推理性能,但这些方法通常依赖重复生成或外部验证。为解决这一局限,我们提出CritICL,一种在保持高效率的同时提升推理能力的新型推理时框架。我们的核心洞见是,同一模型系列内不同规模的LLM失败模式呈现结构化模式,CritICL不将失败视为不良输出,而是将其用作指导来源。具体而言,我们利用从较弱模型得到的失败模式,通过基于批判的上下文示例将其融入推理。我们提出两种变体:CritICL-dynamic,自适应预测输入特定的失败模式并检索批判;CritICL-static,使用全局失败模式轮廓提供稳定指导。实验结果显示,CritICL始终优于标准上下文学习,且性能与测试时缩放方法相当或更优,同时需要显著更少的生成次数和更低的token成本。代码可在以下网址获取:this https URL
英文摘要
Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs). However, these methods typically rely on repeated generation or external verification. To address this limitation, we introduce CritICL, a novel inference-time framework that improves reasoning while maintaining high efficiency. Our key insight is that LLM failure modes exhibit structured patterns across model scales within the same family. Instead of treating failures as undesirable outputs, CritICL leverages them as a source of guidance. Specifically, we utilize failure modes derived from weaker models and incorporate them into inference through critique-based in-context examples. We propose two variants: CritICL-dynamic, which adaptively predicts input-specific failure modes and retrieves critiques, and CritICL-static, which uses a global failure mode profile to provide stable guidance. Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost. Code available at: https://github.com/umwyf/CRITICL
Commentshttps://github.com/umwyf/CRITICL