发表机构
Rakuten Group, Inc.(乐天集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有迭代优化方法忽略奖励衰减导致过度利用的问题,提出结合奖励衰减建模与EM算法的联合LinUCB算法,在Sentiment Reversal和GSM8K基准上显著优于强基线。
AI 中文摘要
迭代优化已显著提升了大语言模型(LLM)的性能;然而,从基于反馈的Self-Refine到传统老虎机方法的现有方法,往往依赖静态选项或忽略饱和效应。这种忽略导致过度利用,即持续使用相同的提示或臂会随时间导致奖励递减。为应对这一挑战,我们提出了一种明确纳入奖励衰减建模的新型上下文老虎机算法。利用期望最大化(EM)算法,我们的方法同时估计特定臂参数和衰减参数。此外,通过将提示嵌入为臂,我们促进了臂值的联合学习,这与传统的不相交线性上置信界(LinUCB)框架形成区别。在情感反转(Sentiment Reversal)和GSM8K基准上的实验结果表明,我们的方法相较于强基线取得了显著的性能提升。最后,我们的 ablation 研究证实,在老虎机框架内整合奖励衰减建模对于缓解过度利用和优化迭代优化过程至关重要。
英文摘要
Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-based Self-Refine to traditional bandit approaches often rely on static options or overlook the saturation effect. This neglect leads to over-exploitation, where the continuous use of identical prompts or arms results in diminishing rewards over time. To address this challenge, we propose a novel contextual bandit algorithm that explicitly incorporates reward decay modeling. Utilizing an Expectation-Maximization (EM) algorithm, our method simultaneously estimates both arm-specific and decay parameters. Furthermore, by embedding prompts as arms, we facilitate the joint learning of arm values, distinguishing our approach from the traditional disjoint Linear Upper Confidence Bound (LinUCB) framework. Experimental results on Sentiment Reversal and GSM8K benchmarks demonstrate that our method achieves significant performance gains over strong baselines. Finally, our ablation study confirms that the integration of reward decay modeling within the bandit framework is crucial for mitigating over-exploitation and optimizing the iterative refinement process.