面向未知奖励的推理时对齐理论
Towards a theory of inference-time alignment with unknown rewards
浏览论文内容
中文总结 AI 辅助
该研究将推理时对齐形式化为弱到强学习问题,引入对齐维度表征其可学习性,证明其有限性是可对齐学习的充要条件,为对齐理论构建提供了思路。
中文摘要 AI 辅助
生成式模型对齐已受到广泛关注,在监督微调与推理时计算方面取得了显著进展,但从统计学习视角对齐仍未得到充分理解。我们将推理时对齐形式化为弱到强学习问题,其中假设参考策略(弱学习器)性能较好,目标是生成强学习器,使其在测试时以任意高概率预测优质响应。我们的问题是从头学习的——所有内容均从数据中学习,而非假设可获取优质奖励估计,因此与现有推理时对齐理论不同。我们的模型与 arXiv:2510.15464 的近期工作相似,该工作中每个提示可能存在多个优质响应。我们的对齐可学习性定义遵循 PAC 学习原理。我们引入了奖励类别的一种新型组合维度,将其命名为对齐维度,并证明其可完全表征对齐可学习性——当且仅当奖励类别的对齐维度有限时,该奖励类别是可对齐学习的。我们学习过程的核心是调用普通的单包含图算法,对所有满足互不包含的标签集对进行锦标赛。我们认为,我们的结果可能为建立对齐的完整理论理解提供思路。
英文摘要
Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation. Yet, alignment has remained poorly understood from a statistical learning perspective. We formulate inference-time alignment as a weak-to-strong learning problem, where a reference policy (weak model) is assumed to be fairly good and the goal is to produce a strong model that predicts a good response at test time with arbitrarily high probability. Our problem is formulated as learning from scratch --- everything is learned from data rather than assuming access to a good reward estimate, and thus differs from the existing inference-time alignment theory. Our framework shares similarity to the recent work of Joshi et al., (arXiv:2510.15464), where for each prompt, there could be multiple good responses. Our definition of the alignment learnability follows the standard PAC learning principle. We introduce a novel combinatorial dimension of the reward class which we call the alignment dimension, and show that it completely characterizes the alignment learnability --- a reward class is alignment learnable if and only if its alignment dimension is finite. The core of our learning procedure works by learning a pairwise comparator and then running a tournament over candidate responses. We believe that our results might shed light toward establishing a complete theoretical understanding of alignment.
发表机构
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。