arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10848cs.LG

大语言模型的摊销式离线策略评估

Amortized Off-Policy Evaluation for LLMs

Younwoo Choi, Leo Feng, Vincent Liu, Haanvid Lee

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM持续部署中的策略与奖励偏移问题,提出PFN-OPE摊销式离线策略评估方法,在两个数据集上较基线误差降低2.0至9.3倍。

中文摘要 AI 辅助

准确的评估是选择部署哪款大语言模型(LLM)的核心,但在真实流量中测试候选模型会让真实用户接触未经过充分验证的模型。因此,团队会在离线环境下,利用已部署模型产生的数据评估候选模型,这就是离线策略评估(OPE),它面临两种分布偏移:一是模型在后续训练中更新后,其响应与记录的响应产生差异(策略偏移);二是用于评判模型的奖励定义会随业务需求变化(奖励偏移)。经典OPE方法不适用于这种持续部署场景,因为它们是针对特定任务定义的,需要针对每个新的记录数据集或奖励定义从头拟合。为解决该问题,我们提出PFN-OPE,这是一种基于先验数据拟合的网络,可在上下文博弈任务分布上摊销OPE。我们用由多个奖励函数评分的LLM响应池构建的任务对其进行一次预训练,该任务中存在上述两种偏移。在测试时,它会将一个记录数据集和每个提示对应的一个采样目标响应映射为一个价值估计,仅需一次前向传播,无需针对特定任务拟合。在HelpSteer2和UltraFeedback数据集上,针对Qwen、Llama和Gemma策略,在奖励偏移场景的所有测试配置中,PFN-OPE的误差比最佳基线低2.0至9.3倍。

英文摘要

Accurate evaluation is central to selecting which LLM to deploy, yet testing a candidate on live traffic exposes real users to an unvetted model. Teams therefore evaluate candidates offline, on data produced by already-deployed models. This is off-policy evaluation (OPE), and it faces two distribution shifts: as a model is updated in post-training, its responses diverge from the logged ones (policy shift), and the reward definition under which it is judged changes with business requirements (reward shift). Classical OPE methods are ill-suited to this continual-deployment setting because they are defined per task and require fitting from scratch on every new logged dataset or reward definition. To address this, we propose PFN-OPE, a prior-data fitted network that amortizes OPE across a distribution of contextual-bandit tasks. We pretrain it once on tasks constructed from a pool of LLM responses scored by several reward functions, in which both shifts occur. At test-time it maps a logged dataset and one sampled target response per prompt to a value estimate in a single forward pass, with no per-task fitting. On HelpSteer2 and UltraFeedback with Qwen, Llama, and Gemma policies, PFN-OPE achieves 2.0 to 9.3 times lower error than the best baselines across all tested configurations in the reward-shifted settings.

发表机构

  • RBC Borealis
  • University of Toronto(多伦多大学)
  • Vector Institute(向量研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑