arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05820cs.LGstat.ML

有限反馈下基于LLM专家的在线学习

Online Learning with LLM Experts from Limited Feedback

  • Virginia Tech(弗吉尼亚理工大学)
  • Adobe Research(Adobe 研究院)

机构由 AI 辅助整理,请以论文原文为准。

Wang Wei, Soumyabrata Pal, Koyel Mukherjee, Franck Dernoncourt, Ryan A. Rossi, Branislav Kveton, Hoda Eldardiry

AI总结:

本文研究有限反馈下将提示路由至LLM专家的在线学习问题,提出老虎机算法实现次线性遗憾,并在实验中验证了高效学习高质量路由策略的能力。

AI中文摘要:

我们研究在有限反馈的在线环境中,将提示自适应路由到大语言模型(LLM)专家以最大化响应质量的问题。我们将其形式化为一个多臂老虎机问题,其中$K$个动作代表专家,$d$个特征编码提示,时间范围为$T$轮。我们提出了能够策略性地选择并观察奖励以最小化遗憾的算法。在完全信息设置下,我们实现了$\tilde{O}(d T / \sqrt{m})$的遗憾界,而在老虎机设置下,我们实现了$\tilde{O}(d T \sqrt{K / m})$的遗憾界,其中$m \ll T$是反馈预算。我们的实验表明,我们能够从有限反馈中高效地学习跨不同LLM的高质量路由策略。

英文摘要:

We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online setting with limited feedback. We formulate it as a bandit problem with $K$ actions that represent experts and $d$ features that encode prompts, over a horizon of $T$ rounds. We propose algorithms that strategically select and observe rewards to minimize regret. In the full-information setting, we achieve a regret of $\tilde{O}(d T / \sqrt{m})$, while in the bandit setting we achieve $\tilde{O}(d T \sqrt{K / m})$, where $m \ll T$ is a budget on feedback. Our experiments show that we efficiently learn high-quality routing strategies across diverse LLMs from limited feedback.

补充信息

↑