发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型少样本归纳推理能力不足的问题,提出沙漏推理方法,通过在推理阶段强制实现严格上下文隔离,构建符号编码器 - 解码器。实验表明该方法在多基准测试中显著提升准确率,证实收益源于阶段隔离和初始归纳质量。
AI 中文摘要
自我优化往往无法增强大语言模型中的少样本归纳推理能力。仅提示模型明确陈述其推断规则作用不大。关键在于推理阶段之间在结构上强制隔离,使信息只能以压缩符号状态在它们之间传递。我们引入了沙漏推理,它在推理阶段之间强制实现严格的上下文隔离。冻结的大语言模型充当元构造器,为每个任务构建一个符号编码器 - 解码器:归纳模块将支持示例压缩为模式$\phi$(编码器)和临时支架$z$;演绎模块从这些中推导规则$T$(解码器)并丢弃$z$;实现器将$(\phi, T)$编译为工件;错误驱动的优化器修改$(\phi, T)$并从头重新生成工件。只有$(\phi, T)$跨越阶段边界,所以所有优化都锚定在规则上。我们使用GPT - 5.5和Gemini 3.1 Pro在跨越视觉抽象、硬件合成和文本规则归纳的三个基准上评估沙漏推理。在ARC - AGI - 2上,它比迭代优化基线将最佳5次准确率提高了多达14分。在ChipBench上,使用GPT - 5.5时,它几乎使Verilog合成准确率翻倍,从31%提高到58%。BBEH - Linguini借鉴了国际语言学奥林匹克竞赛的谜题,在这种设置下先前的工作表明明确的语言表达会损害性能。沙漏推理减轻了这种趋势,在Gemini 3.1 Pro上,它完全扭转了这种影响。消融实验证实这些收益来自阶段之间的隔离和初始归纳的质量,而不是提示措辞或使用的特定符号形式。是信息在推理过程中的流动方式,而非用于表达它的语言,驱动了冻结大语言模型中的归纳推理。
英文摘要
Self-refinement often fails to strengthen few-shot inductive reasoning in large language models. Prompting a model to explicitly state its inferred rule does little on its own. What actually matters is a structurally enforced isolation between reasoning stages, so that information can only pass between them as a compressed symbolic state. We introduce \textbf{Hourglass reasoning}, which enforces strict context isolation between reasoning stages. The frozen LLM acts as a meta-constructor, building for each task a symbolic encoder--decoder: an Induction module compresses the support examples into a schema $ϕ$ (encoder) and a transient scaffold $z$; a Deduction module derives rule $T$ (decoder) from these and discards $z$; an Implementer compiles $(ϕ, T)$ into artifacts; an error-driven Refiner revises $(ϕ, T)$ and regenerates artifacts from scratch. Only $(ϕ, T)$ crosses stage boundaries, so all refinement stays anchored to the rule. We evaluate Hourglass across three benchmarks spanning visual abstraction, hardware synthesis, and textual rule induction, using GPT-5.5 and Gemini 3.1 Pro. On ARC-AGI-2, it raises best-of-5 accuracy by up to 14 points over an iterative-refinement baseline. On ChipBench, it nearly doubles Verilog synthesis accuracy with GPT-5.5, from 31\% to 58\%. BBEH-Linguini draws on puzzles from the International Linguistics Olympiad, a setting where prior work has shown that explicit verbalization can hurt performance. Hourglass mitigates this tendency, and on Gemini 3.1 Pro, it reverses the effect entirely. Ablations confirm that these gains come from the isolation between stages and the quality of the initial induction, not from prompt wording or the particular symbolic form used. It is how information flows through the reasoning process, rather than the language used to express it, that drives inductive reasoning in frozen LLMs.