快速A/B/n测试:通过树耦合反馈共享实现精确多策略比较
Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing
浏览论文内容
中文总结 AI 辅助
该研究提出树耦合A/B测试(TCAB)方法,通过树耦合反馈共享实现多策略精确比较,将奖励查询数从JT降至T+o(T),实验显示其在成本-精度上有显著提升。
中文摘要 AI 辅助
在线平台越来越多地比较多种自适应决策策略,包括排序系统、推荐算法、定价规则和语言模型智能体,而每个带有奖励的交互可能成本高昂或存在风险。直接的A/B/n设计为每个J个策略提供各自的T步轨迹,因此需要使用JT个结果。我们提出树耦合A/B测试(TCAB),这是一种适用于任意依赖历史的上下文博弈策略的精确反馈共享设计。在每一轮,一棵可预测的树连接当前的策略历史;每个父子上下文-动作律被最大程度地耦合,且在匹配树边的每个组件内共享一个奖励。即使策略被故意依赖,每个策略仍保留其独立的有限步轨迹律。若D_{e,t}记录轮t时树边e上的不匹配,奖励查询数满足路径恒等式N(T)=T+∑_{t,e}D_{e,t},因此等于T加上期望中累积的树边总变差。该成本在所选树的精确边局部设计中是条件最优的,且当前轮的最小生成树在树设计中是近视最优的。对于固定的J,每个策略的次线性伪遗憾和神谕动作的几乎必然唯一性意味着E[N(T)]=T+o(T),而独立运行则为JT。我们还获得了成对策略对比的有限样本方差界。在奖励模型评估、多选项语言模型评估和自适应搜索策略上的实验表明,在成本-精度前沿有显著提升。
英文摘要
Online platforms increasingly compare many adaptive decision policies---ranking systems, recommendation algorithms, pricing rules, and language-model agents---while each reward-bearing interaction can be costly or risky. A direct A/B/n design gives each of $J\ge 2$ policies its own horizon-$T$ trajectory and therefore uses $JT$ outcomes. We introduce Tree-Coupled A/B Testing (\TCAB), an exact feedback-sharing design for arbitrary history-dependent contextual-bandit policies. At each round, a predictable tree connects the current policy histories; every parent--child context--action law is maximally coupled, and one reward is shared within each component of matched tree edges. Every policy retains exactly its standalone finite-horizon trajectory law, even though the policies are deliberately dependent. If $D_{e,t}$ records a mismatch on tree edge $e$ at round $t$, the number of reward queries satisfies the pathwise identity $N(T)=T+\sum_{t,e}D_{e,t}$ and hence equals $T$ plus cumulative tree-edge total variation in expectation. This cost is conditionally optimal among exact edge-local designs on the selected tree, and a current-round minimum-spanning tree is myopically optimal among tree designs. For fixed $J$, sublinear pseudo-regret of every policy and almost-sure uniqueness of the oracle action imply $\mathbb{E}[N(T)]=T+o(T)$, versus $JT$ for independent runs. We also obtain finite-sample variance bounds for pairwise policy contrasts. Experiments on reward-model evaluation, multiple-choice language-model evaluation, and adaptive search policies demonstrate substantial improvements in the cost--precision frontier.
发表机构
- Courant Institute of Mathematical Sciences, New York University(纽约大学柯朗数学科学研究所)
机构由 AI 辅助整理,请以论文原文为准。