arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于Plackett-Luce最佳K目标的秩条件样本重用

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective

Melveena Jolly, Midhun Xavier

arXiv 2607.11146首次发表:更新:

发表机构

Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究耦合目标\(J_K^{WOR}\),现有样本重用方法在此有偏差。提出用秩条件霍维茨 - 汤普森估计为\(J_K^{WOR}\)子集总量构建无偏估计器,含动态规划简化计算,还介绍了相关性质及未解决问题。

AI 中文摘要

我们研究了耦合目标\(J_K^{WOR}=E_{S \sim PL - WOR_K}[max_{i \in S} R_i]\):即无放回抽取大小为\(K\)的Plackett-Luce样本的预期最大奖励,以及Gumbel-Top-K/随机束搜索解码法则。这个估计量与现有样本重用的Max@K估计器所针对的传统独立同分布目标\(J_K^{iid}=E[max_{i \leq K} R_i]\)不同,并且在耦合采样器下重用它们的独立同分布权重会产生偏差(一个封闭形式的三项实例给出\(E[g_{iid}] = \frac{4}{5} grad J_K^{WOR}\);耦合采样器下的pass@K是二元奖励的特殊情况)。通用联合分数REINFORCE对\(J_K^{WOR}\)已经是无偏的;它所缺乏的是样本重用。我们的贡献是为\(J_K^{WOR}\)子集总量实例化标准的秩条件霍维茨 - 汤普森估计:从一个Gumbel-Top-n池(\(n>K\))及其观察到的优先级阈值,我们构建了一个估计器,该估计器重用所有\(C(n,K)\)个嵌入的\(K\)子集,通过无偏精确得分函数替代梯度实现无偏,加上一个奖励排序的Max特定动态规划,将\(C(n,K)\)项子集和(具有\(K!\)成本集概率)精确地简化为一维积分。固定\(Q\)求积评估的算术成本为\(O(n log n + nKQ)\),并且在数值上而非代数上是精确的;没有认证的\(\epsilon\)近似率。当\(n \geq 2K\)时,每个非零的\(K\)阶霍维茨 - 汤普森项都具有有限的二阶矩;在相同假设下,只要\(n \geq 2K\),完整的替代梯度就具有有限的二阶矩(那里的尖锐度是开放的)。当\(K = 1\)时,该构造恢复经典优先级采样。所有数量仅需要\(n + 1\)个抽取项目概率的值和可微计算图,因此涵盖了通过精确SBS采样的有限结构化序列策略。一个认证的有限\(Q\)求积界和可数无限支持仍然是开放的。验证代码作为辅助文件包含在内(这里的代码相关表述是按原文翻译,实际可能需要根据具体情况调整理解)。

英文摘要

We study the coupled objective J_K^WOR = E_{S ~ PL-WOR_K}[max_{i in S} R_i]: the expected maximum reward of a size-K Plackett-Luce draw without replacement, the law of Gumbel-Top-K / Stochastic Beam Search decoding. This estimand differs from the conventional i.i.d. objective J_K^iid = E[max_{i<=K} R_i] targeted by existing sample-reuse Max@K estimators, and reusing their i.i.d. weights under the coupled sampler is provably biased (a closed-form three-item instance gives E[g_iid] = (4/5) grad J_K^WOR exactly; pass@K under the coupled sampler is the binary-reward special case). Generic joint-score REINFORCE is already unbiased for J_K^WOR; what it lacks is sample reuse. Our contribution is to instantiate standard rank-conditioned Horvitz-Thompson estimation for the J_K^WOR subset total: from one Gumbel-Top-n pool (n>K) and its observed priority threshold we build an estimator that reuses all C(n,K) embedded K-subsets, unbiased with an unbiased exact score-function surrogate gradient, plus a reward-sorted Max-specific dynamic program that collapses the C(n,K)-term subset sum (with K!-cost set probabilities) exactly to a one-dimensional integral. A fixed-Q quadrature evaluation costs O(n log n + nKQ) arithmetic and is numerically, not algebraically, exact; no epsilon-approximation rate is certified. Each nonzero degree-K Horvitz-Thompson term has finite second moment exactly when n >= 2K; under the same assumptions the full surrogate gradient has finite second moment whenever n >= 2K (sharpness there is open). At K=1 the construction recovers classical priority sampling. All quantities require only the values and differentiable computation graphs of the n+1 drawn items' probabilities, so finite structured sequence policies sampled by exact SBS are covered. A certified finite-Q quadrature bound and countably infinite support remain open. Validation code is included as ancillary files.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑