发表机构
Technion – Israel Institute of Technology(以色列理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对生成建模中条件采样的挑战,提出基于偏好投票的精确MH采样器Pref-MH,利用成对比较反馈实现最优条件采样,在多模态生成任务中验证了其实用性。
AI 中文摘要
从具有期望语义属性的分布中采样是现代生成建模中的新兴挑战。Metropolis-Hastings(MH)为条件采样提供了原则性途径,但需要精确的逐点目标密度评估,而在生成设置中无法获取此类评估。与此同时,人类或模型“评判者”的成对比较极易获取,且已在多种应用中被证明具有重要价值。我们提出Pref-MH,一种仅使用随机二元成对比较的、针对评判者诱导的条件分布的通用精确MH采样器。我们的核心发现是,MH的未归一化密度比率与Bradley-Terry(BT)选择模型的偏好比率相匹配。核心挑战在于,MH需要精确的比率计算,而BT评判者仅提供采样的二元反馈。为此,我们开发了一种有效的接受/拒绝规则,其生成的马尔可夫链可被证明收敛至目标分布。我们进一步证明,对于固定的建议核和预算,Pref-MH在这类精确可逆接受规则中具有Peskun-Tierney意义上的最优性。在文本生成和分子设计(使用LLM评判者)以及图像生成(使用VLM评判者)上的实验表明,当相对容易获得比较反馈时,Pref-MH提供了一种实用且灵活的条件采样方法。
英文摘要
Sampling from distributions conditioned on desired semantic properties is an emerging challenge in modern generative modeling. Metropolis-Hastings (MH) provides a principled route to conditional sampling, but requires access to exact pointwise target-density evaluations, which are not available in generative settings. Meanwhile, pairwise comparisons by humans or model "judge" are highly accessible and have proved valuable across diverse applications. We introduce Pref-MH, a general exact MH sampler for judge-induced conditional distributions using only stochastic binary pairwise comparisons. Our key observation is that the MH unnormalized density ratio matches the preference odds of the Bradley-Terry (BT) choice model. The central challenge is that while MH requires precise ratio computation, BT judges provide only sampled binary feedback. To this end, we develop a valid accept/reject rule whose resulting Markov chain provably converges to the target distribution. We further show that, for a fixed proposal kernel and budget, Pref-MH is optimal in the Peskun-Tierney sense among this class of exact reversible acceptance rules. Experiments on text generation and molecular design with LLM judges, as well as image generation with VLM judges, demonstrate that Pref-MH provides a practical and flexible approach to conditional sampling when comparative feedback is relatively easy to obtain.