通过拉格朗日奖励增强实现安全推理时对齐
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
浏览论文内容
中文总结 AI 辅助
研究提出拉格朗日奖励增强框架,从含奖励与成本模型的约束目标出发,经对偶化将问题转化为一维凸问题,校准对偶变量获增强奖励,提升推理时对齐的有益性与无害性权衡。
中文摘要 AI 辅助
推理时对齐利用辅助奖励信号在解码时引导冻结语言模型,避免重复权重更新成本。现有方法通常优化单个标量分数,安全约束易被忽略或需手动调整惩罚。我们提出拉格朗日奖励增强(LARA),一种安全约束下的通用推理时对齐框架……
英文摘要
Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize a single scalar score, so explicit safety constraints must either be ignored or encoded through manually tuned penalties. We propose Lagrangian Reward Augmentation (LARA), a general inference-time alignment framework under safety constraints. Starting from a KL-regularized constrained objective with a reward model and a cost model, LARA dualizes the constraint and reduces the optimization problem to a one-dimensional convex problem over a nonnegative dual variable. Estimated on a small calibration set, this dual variable defines an augmented reward that can be used as a drop-in scoring signal within existing inference-time alignment methods. For sequence-level sampling methods, such as Best-of-N reranking, the calibrated dual variable corresponds to the solution of the expected-cost constrained problem. For token-level reward-guided decoding methods, the same construction yields a principled dual-calibrated heuristic rather than an exact constrained-policy guarantee. We evaluate LARA on both sequence-level and token-level inference-time alignment methods, and find that LARA improves the helpfulness-harmlessness tradeoff, with Best-of-N achieving the best performance among inference-time methods, approaching finetuning-based direct alignment baselines.
发表机构
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
- Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。