Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
Re-FORC:用于高效思维链推理的自适应奖励预测
机构 * AWS Agentic AI(AWS智能代理部门) ; Carnegie Mellon University(卡内基梅隆大学)
AI总结 研究提出Re-FORC自适应奖励预测方法,通过在推理模型上训练轻量级适配器,实现早期停止无前景推理链、优化模型和思维长度选择以及自适应测试时缩放,有效提升推理效率和准确率。
Comments Accepted at ICML 2026; previously accepted as a non-archival paper at the Efficient Reasoning Workshop at NeurIPS 2025