发表机构
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ReHoPER,一种仅推理的零样本方法,通过滚动规划中间问题并回答来增强LLM推理,在组合推理基准上优于强基线。
AI 中文摘要
我们提出ReHoPER,一种仅推理、零样本的方法,通过在多条路径上生成并回答中间问题,在最终答案之前提升大型语言模型的推理能力。它迭代地规划一个候选中间问题的时域,选择一个进行回答,并从更新后的历史中重新规划。ReHoPER是任务无关的,跨数据集和模型使用相同的通用指令,无需标注数据或特定任务的提示设计。在多个数据集上,包括iLLC(一个新的用于组合推理的受控基准),ReHoPER优于强基线,在最组合的设置中取得最大提升。我们的实现和iLLC生成器公开可用,以支持未来工作。
英文摘要
We propose ReHoPER, an inference-only, zero-shot method that improves large language models' reasoning by generating and answering intermediate questions along multiple paths before the final answer. It iteratively plans a horizon of candidate intermediate questions, selects one to answer, and replans from the updated history. ReHoPER is task-agnostic, using the same generic instructions across datasets and models without labeled data or task-specific prompt design. Across multiple datasets, including iLLC, a new controlled benchmark for compositional reasoning, ReHoPER outperforms strong baselines, with the largest gains in the most compositional settings. Our implementation and the iLLC generator are publicly available to support future work.