发表机构
University of Birmingham(伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出STRETCH统一框架,通过动态伸展区机制对齐问题难度与模型能力,采用支架者与学习者双循环协同进化,在谈判和运筹学基准上超越强基线,实现渐进式推理增长。
AI 中文摘要
大型语言模型(LLMs)在自我改进训练中常因能力停滞而受限,因为固定的难度水平无法适应其不断发展的熟练度。为解决此问题,我们提出STRETCH(通过针对性挑战进行自教推理进化),一个受认知支架理论启发的统一框架。STRETCH引入动态伸展区机制,持续将问题难度与模型的解题能力对齐。在单一参数空间内,模型在生成自适应、突破边界挑战的支架者与通过强化学习优化其解题轨迹的学习者之间交替。这种双循环协同进化有效稳定训练,缓解奖励黑客问题,并促进渐进式推理增长。在谈判和运筹学基准上的实验表明,STRETCH始终优于强提示和领域特定基线。进一步的支架者配置分析显示,动态难度对齐对于持续能力提升和同步推理进化至关重要。
英文摘要
Large language models (LLMs) often suffer from capability stagnation in self-improvement training because fixed difficulty levels fail to adapt to their evolving proficiency. To address this issue, we propose STRETCH (Self-Taught Reasoning Evolution via Targeted CHallenge), a unified framework inspired by cognitive scaffolding theory. STRETCH introduces a dynamic Stretch Zone mechanism that continuously aligns question difficulty with the model's solving capability. Within a single parameter space, the model alternates between a Scaffolder that generates adaptive, boundary-pushing challenges and a Learner that that optimizes its solving trajectories through reinforcement learning. This dual-loop co-evolution effectively stabilizes training, mitigates reward hacking and promote progressive reasoning growth. Experiments on both negotiation and operation research benchmarks demonstrate that STRETCH consistently outperforms strong prompting and domain-specific baselines. Further scaffolder configuration analysis shows that dynamic difficulty alignment is critical for sustained capability improvement and synchronized reasoning evolution.