arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

均匀竞赛:无参数近似拒绝采样

Uniform Race: Parameter-Free Approximate Rejection Sampling

Seiyun Shin, Juhyeong Pang, Kwang-Sung Jun

arXiv 2609.34639首次发表:更新:

发表机构

Pohang University of Science and Technology; University of Wisconsin–Madison(浦项科技大学; 威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对近似拒绝采样需预设阈值的局限,提出无参数算法均匀竞赛(UR),通过重要性权重除以均匀随机变量选取最大分数样本,在任意预算下同时达到所有阈值的最优误差界,且优于现有基线,并在LLM数学推理任务中验证有效性。

AI 中文摘要

我们研究近似采样问题:给定来自提议分布 $\mu$ 的 $N$ 个独立样本,目标是选择一个样本,使其分布接近仅由归一化常数确定的目标分布 $\pi$。Block 和 Polyanskiy (2023) 针对近似拒绝采样(RS)提供了有限预算误差界,该误差界是接受阈值 $M$ 的函数。然而,给出最小误差界的阈值 $M$ 依赖于 $\pi$ 和 $\mu$ 的性质,而这些性质通常无法从观测样本中获得。这引出一个自然问题:能否在不将 $M$ 作为输入的情况下达到最佳的 RS 保证?我们给出肯定回答,提出一种基于重要性权重的无参数采样算法,称为均匀竞赛(UR),其中重要性权重是目标概率与提议概率(或密度)的比值。该算法将每个观测权重除以一个独立的均匀随机变量以形成分数,并返回具有最大分数的候选样本。对于每个预算 $N$,其总变差误差同时满足每个固定阈值 $M$ 的 RS 上界,从而事后达到此类界中的最优值。我们还刻画了在最大分数条件下的输出分布,确定了何时该分布恰好为目标分布 $\pi$。均匀竞赛的总变差误差不大于由 Rohatgi 等人 (2025) 推导的自然预算校准 RS 以及采样重要性重采样(SIR)的误差。特别地,我们展示了在某些实例中,UR 的误差在 $N$ 上比两种基线方法呈指数级更小。此外,我们建立了条件,在这些条件下,对于每个 $\pi$ 和 $\mu$ 达到此 RS 保证会唯一确定选择概率为 UR 的选择概率。最后,在 LLM 数学推理任务上的测试时缩放实验证实了理论比较,并表明 UR 在无需阈值选择的情况下仍能保持与真实准确率相当的竞争力。

英文摘要

We study approximate sampling: given $N$ independent samples from a proposal distribution $μ$, the goal is to select one whose distribution is close to a target $π$ specified only up to a normalizing constant. Block and Polyanskiy (2023) provide finite budget error bounds for approximate rejection sampling (RS) as a function of the acceptance threshold $M$. The threshold $M$ giving the smallest bound, however, depends on properties of $(π,μ)$ that are typically unavailable from the observed sample. This raises a natural question: Can one attain the best RS guarantee without taking $M$ as input? We answer affirmatively by proposing a parameter-free sampling algorithm called uniform race (UR), based on importance weights, which are ratios of target to proposal probabilities (or densities). It divides each observed weight by an independent uniform random variable to form a score and returns the candidate with the largest score. For every budget $N$, its total variation error satisfies the RS upper bound for every fixed threshold $M$ simultaneously, thereby achieving the best such bound in hindsight. We also characterize its output distribution conditional on the largest score, identifying when it is exactly the target $π$. Uniform race has no larger total variation error than a natural budget-calibrated RS derived from Rohatgi et al. (2025) and sampling importance resampling (SIR). In particular, we exhibit instances where UR's error is exponentially smaller in $N$ than that of either baseline. Furthermore, we establish conditions under which attaining this RS guarantee for every $(π,μ)$ uniquely determines the selection probabilities as those of UR. Finally, test-time scaling experiments on LLM math-reasoning tasks corroborate the theoretical comparisons and demonstrate that UR remains competitive in ground-truth accuracy without requiring threshold selection.

Comments45 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑