自优化基础模型的最优停止
Optimal Stopping of Self-Refining Foundation Models
浏览论文内容
中文总结 AI 辅助
本文将基础模型的自优化过程形式化为最优停止问题,推导可高效计算的最优停止策略,实验表明该策略在编码基准上比现有方法成本效益更高。
中文摘要 AI 辅助
基础模型可通过外部反馈驱动的自优化过程提升输出质量,该过程中模型处于迭代循环:生成输出、接收验证器反馈、通过上下文学习优化响应。本文将此过程形式化为最优停止问题,根据预期改进与成本的关系决定优化迭代次数,推导最优停止策略并证明其可通过随机近似高效计算。实验将该策略应用于基础模型的编码基准,结果显示其停止策略比现有工作的停止策略成本效益显著更高。
英文摘要
Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is embedded in an iterative loop where it generates outputs, receives feedback from verifiers, and refines its responses through in-context learning. Following a novel approach, we formalize this process as an optimal stopping problem where the number of refinement iterations is decided based on expected improvement relative to cost. We derive optimal stopping policies and show that they can be efficiently computed through stochastic approximation. To evaluate our approach experimentally, we apply it to a coding benchmark for foundation models. The empirical results show that our stopping policies are significantly more cost-efficient than stopping policies proposed in prior work.
发表机构
- University of Melbourne(墨尔本大学)
- Imperial College London(伦敦帝国学院)
机构由 AI 辅助整理,请以论文原文为准。