发表机构
The Pennsylvania State University(宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究用于大语言模型路由的相关感知上下文博弈问题,提出耦合奖励混合与解耦预测混合两种利用替代奖励的算法,经理论分析和实验评估,相比基线方法提高了样本效率和精度 - 成本权衡。
AI 中文摘要
我们研究了具有相关臂且能获取由机器学习模型产生的替代奖励信号的上下文博弈问题,这是受大语言模型路由等应用的推动。与仅依赖博弈反馈并假设臂间条件独立的经典上下文博弈不同,我们的设置允许上下文相关的臂间相关性以及可能有噪声或错误指定的辅助奖励信息。我们提出了通过两种互补设计利用此类替代奖励的算法。一种耦合奖励混合方法在替代信号可靠时将真实奖励和替代奖励合并以加速学习,而一种解耦预测混合方法为博弈反馈和替代奖励维持单独的估计器并自适应地组合它们的预测。这种解耦在最坏情况下能抵御替代错误指定,恢复与仅奖励博弈方法相当的遗憾保证,而当替代预测足够有信息时能实现更好的遗憾。我们为两种方法提供了理论遗憾分析,并在不同精度与成本权衡下的大语言模型路由基准上对它们进行评估。结果表明与标准上下文博弈基线和强大的静态路由方法相比,样本效率提高且始终具有更好的精度 - 成本权衡。
英文摘要
We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large language model (LLM) routing. Unlike classical contextual bandits that rely solely on bandit feedback and assume conditional independence across arms, our setting allows context-dependent inter-arm correlations and auxiliary reward information that may be noisy or misspecified. We propose algorithms that leverage such surrogate rewards through two complementary designs. A coupled reward-mixing approach pools true and surrogate rewards to accelerate learning when surrogate signals are reliable, while a decoupled prediction-mixing approach maintains separate estimators for bandit feedback and surrogate rewards and adaptively combines their predictions. This decoupling yields robustness to surrogate misspecification, recovering regret guarantees comparable to reward-only bandit methods in the worst case, while achieving improved regret when surrogate predictions are sufficiently informative. We provide theoretical regret analyses for both approaches and evaluate them on LLM routing benchmarks under varying accuracy versus cost trade-offs. The results demonstrate improved sample efficiency and consistently better accuracy-cost trade-offs compared to standard contextual bandit baselines and strong static routing methods.