arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向大语言模型的无训练推理时自我反思与成本受限早停机制

Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models

Wei Yu, Suxing Liu, Minjie Yu, Jiahao Wang, Zhijian Zheng, Haocheng Deng, Bing Li

arXiv 2608.18884首次发表:更新:

发表机构

School of Digital Arts, Jiangxi Arts & Ceramics Technology Institute; School of Computing, Universiti Sains Malaysia(江西工艺美术职业技术学院数字艺术学院; 马来西亚理科大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出无训练的推理时协议EvoResearcher,通过成本受限的自我反思实现大语言模型早停,在多个推理基准上验证其成本受限自我验证的价值。

AI 中文摘要

针对推理类大语言模型(如GRPO)的强化学习训练成本高昂,且需要可控环境,所有贡献均需投入完整训练流程。本文提出EvoResearcher,一种无训练的推理时协议,仅在单个冻结的大语言模型主干上添加成本受限的自我反思。该协议遵循“生成→自我批判→修正”的迭代流程,直至达到最大深度D,或批判返回“CONFIRMED”哨兵信号——这是一种隐式早停机制,使模型在严格计算预算下自我验证答案。四个自我反思元奖励组件(正确性、效率、反思深度、工具调用多样性)作为设计原则,以提示级机制实现,因此无需梯度更新即可累积其收益。我们在Big-Bench Hard(100个问题)上验证该协议,并在同一冻结主干上的GSM8K(500个问题)和MATH(500个问题)上建立跨域行为,在Qwen2.5-72B上实现跨模型复现。所有实验均使用纯推理基准;工具调用多样性组件以提示级形式验证,而环境级和多智能体扩展为未来工作的设计蓝图。在干净的BBH上,该协议未使准确率超出95%威尔逊区间;其价值在于成本受限的自我验证,“CONFIRMED”早停机制以约每个问题2.1代的成本,在相同准确率下终止82%-88%的条目。

英文摘要

Reinforcement-learning training of reasoning LLMs (e.g., GRPO) is expensive and requires a controllable environment, committing every contribution to a full training pipeline. We present EvoResearcher, a training-free, inference-time protocol that adds cost-bounded self-reflection to a single frozen LLM backbone. The protocol iterates generate -> self-critique -> revise until a maximum depth D is reached or the critique returns the CONFIRMED sentinel, an implicit early stop that lets the backbone self-verify its answer under a strict compute budget. Four self-reflective meta-reward components (correctness, efficiency, reflection depth, tool-call diversity) act as design principles instantiated as prompt-level mechanisms, so their benefits accrue with zero gradient updates. We validate the protocol on Big-Bench Hard (100 questions) and establish cross-domain behavior on GSM8K (500) and MATH (500) on the same frozen backbone, with cross-model replication on Qwen2.5-72B. All experiments use pure-reasoning benchmarks; the tool-call diversity component is validated in prompt-level form, and the environment-level and multi-agent extensions are design blueprints left to future work. On clean BBH the protocol does not raise accuracy beyond the 95% Wilson interval; its value is cost-bounded self-verification, with the CONFIRMED early stop terminating 82-88% of items at equal accuracy (about 2.1 generations per question).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑