面向离线策略评估的预算约束多源反事实标注
Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation
浏览论文内容
中文总结 AI 辅助
本文针对上下文多臂老虎机离线策略评估的预算约束多源反事实标注问题,构建整数分配模型并提出带动态规划子例程的优化算法,实验显示其可降低均方误差。
中文摘要 AI 辅助
离线策略评估(OPE)从日志数据中估计目标策略的价值,但有限的行为策略覆盖范围可能导致高方差重加权或奖励模型外推。反事实标注可补充未观测动作的证据,然而包括领域专家和大语言模型(LLM)在内的实用标注源可能存在成本高、有偏差或含噪声的问题。本文研究上下文多臂老虎机OPE场景下此类标注的预算约束获取问题。给定各标注源的特定成本与误差特征,我们针对上下文-动作对及标注源构建整数分配问题,以最小化估计器方差中依赖标注方案的部分。我们通过首次标注阈值和局部标注价值区间刻画标注的价值。针对耦合多源问题,我们开发了带动态规划子例程的 majorization-minimization 算法,该算法可单调优化目标函数。在合成临床场景及LLM标注的教育老虎机场景中的实验表明,与无标注相比,本文提出的分配方法分别将固定特征下的均方误差(MSE)降低了20.58%和10.77%。
英文摘要
Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy. We study budgeted acquisition of such annotations for contextual-bandit OPE. Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan. We characterize when annotations are valuable through a first-annotation threshold and local annotation-value regimes. For the coupled multi-source problem, we develop a majorization-minimization algorithm with dynamic-programming subroutines that monotonically improves the objective. Experiments in synthetic clinical and LLM-annotated education bandits show that our allocation method reduces fixed-profile mean squared error (MSE) by 20.58% and 10.77%, respectively, relative to no annotation.
发表机构
- Johns Hopkins Carey Business School(约翰斯·霍普金斯凯瑞商学院)
- International Computer Science Institute(国际计算机科学研究所)
- Pennsylvania State University(宾夕法尼亚州立大学)
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。