arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37209cs.LGstat.ML

逐点还是成对:成对损失何时能帮助奖励学习?

Pointwise or Pairwise: When Do Pairwise Losses Help Reward Learning, Provably?

发表机构韩国科学技术院 · 浦项科技大学 · 三星研究院
查看机构详情
  • KAIST(韩国科学技术院)
  • POSTECH(浦项科技大学)
  • Samsung Research(三星研究院)

机构由 AI 辅助整理,请以论文原文为准。

Junghyun Lee, Minsoo Ha, Sanghwa Kim, Yeongjong Kim, Eunjee Lee, Seiyun Shin, Kwang-Sung Jun

首次发表
浏览论文内容

中文总结 AI 辅助

本文在分组离线上下文老虎机中,证明成对损失(VDR)在有限类中优于逐点损失(VR),但在线性类中两者存在偏差-方差权衡,取决于特征几何和误设程度。

中文摘要 AI 辅助

成对损失越来越多地用于奖励学习,即使在观察到逐点奖励的情况下也是如此,但实证结果好坏参半。成对损失何时以及为何优于逐点损失?我们在分组离线上下文老虎机设置中研究这个问题,该设置允许每个上下文有多个动作,涵盖了多种奖励学习场景。我们比较了值回归(VR),它逐点回归观察到的奖励,与值差回归(VDR),它回归在同一上下文下采样的两个动作之间的奖励差异。我们考虑一个半参数模型,其中平均奖励是可学习的动作相关组件和任意的上下文相关但动作无关的干扰项之和,捕捉特定上下文的干扰。使用统一的局部化分析,我们证明了有限和线性函数类的有限样本回归保证,并将其转化为离线遗憾界。对于有限类,VDR消除了VR界中的误设项,并通过在每个上下文内对动作进行平均来改进一个依赖奖励尺度的误差项,这一好处在相应的VR项中不存在。对于线性类,两种方法都没有统一优势:上下文内差分消除了干扰引起的偏差,但当误设足够低时,相对于使用绝对奖励可能会增加估计方差。这产生了一个特征几何相关的偏差-方差权衡,我们通过数值实验加以证实。

英文摘要

Pairwise losses are increasingly used for reward learning even when pointwise rewards are observed, with mixed empirical results. When and why do pairwise losses outperform pointwise losses? We study this question in a grouped offline contextual-bandit setting allowing multiple actions per context, capturing many reward learning scenarios. We compare Value Regression (VR), which regresses observed rewards pointwise, with Value Difference Regression (VDR), which regresses reward differences between a pair of actions sampled under the same context. We consider a semiparametric model where the mean reward is the sum of a learnable action-dependent component and an arbitrary context-dependent yet action-independent nuisance, capturing context-specific disturbances. Using a unified localized analysis, we prove finite-sample regression guarantees for finite and linear function classes and translate them into offline-regret bounds. For finite classes, VDR eliminates the misspecification term in the VR bound and improves a reward-scale-dependent error term by averaging over actions within each context, a benefit absent from the corresponding VR term. For linear classes, neither method uniformly dominates: within-context differencing removes nuisance-induced bias but may increase estimation variance relative to using absolute rewards when the misspecification is sufficiently low. This yields a feature geometry-dependent bias-variance tradeoff, which we corroborate with numerical experiments.

补充信息

↑