发表机构
University of Southern California; University of California, Berkeley; University of Washington, Seattle(南加州大学; 加州大学伯克利分校; 华盛顿大学西雅图分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对离线策略评估,提出方差最优的顺序标注概率方法,利用有限标注预算实现双稳健估计,在模拟和真实数据上显著降低RMSE。
AI 中文摘要
离线强化学习和离线策略评估基于部署前回顾性收集的数据来评估动态治疗规则。在近期的人工智能应用中,状态和奖励信息被记录为复杂的文本或图像,而诸如“LLM作为评判者”等近期人工智能进展可以以未知偏差对其进行标注。专家标注可能可用,但成本更高。例如,通过廉价但不完美的分类器进行安全分类,与昂贵的专家审查相比。我们展示了如何通过具有缺失奖励的双稳健OPE,将有限的地面真值数据标注预算用于顺序离线策略评估,并优化了方差最优的标注概率,其中目标策略值从标注数据中估计。我们刻画了顺序前向单调标注协议的最优标注概率,并提供了一种可行的批量自适应实现。我们的工作源于与一家无家可归者服务非营利组织的合作,该组织随时间为个体撰写案例记录。我们的方法可用于从案例记录数据中解锁可信推断,并回答新的推断性问题,例如:随时间扩大外展努力如何影响住房申请进展和住房安置改善?在模拟和两个真实数据集(来自非营利组织的案例记录和来自LMArena的人类偏好投票)上,在标注预算为完整标注的40%及以上时,住房安置的RMSE降低了34-65%,住房申请进展的RMSE降低了17-68%,而在LMArena上,在每个预算下RMSE降低了55-62%。
英文摘要
Offline reinforcement learning and off-policy evaluation evaluates dynamic treatment rules based on retrospectively collected data prior to deployment. In recent AI applications, state and reward information is recorded as complex text or image, which recent AI advancements such as LLM-as-a-judge can label with unknown bias. Expert annotation may be available but at a higher cost. For example, safety classification via cheap but imperfect classifiers vs. expensive expert review. We show how a limited budget for ground-truth data-annotation can be used via doubly-robust OPE with missing rewards, and we optimize variance-optimal annotation probabilities for sequential off-policy evaluation, where the target policy value is estimated from annotated data. We characterize the optimal annotation probabilities for sequential forward-monotone annotation protocols, and provide a feasible batch-adaptive implementation. Our work is motivated by a collaboration with a homelessness services nonprofit that writes casenotes for individuals over time. Our method can be used to unlock trustworthy inference from casenote data and answer new inferential questions such as: how does expanding outreach effort over time affect progress towards a housing application and improvement in housing placement? In simulations and on two real datasets - casenotes from the nonprofit and human-preference votes from LMArena - we see reductions in RMSE of 34-65% for housing placement and 17-68% for progress towards a housing application at budgets of 40% of full annotation and above, and by 55-62% at every budget on LMArena.