arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.02781cs.LGcs.AIcs.CL

通过拉格朗日奖励增强实现安全推理时对齐

Safe Inference-Time Alignment via Lagrangian Reward Augmentation

Yaswanth Chittepu, Ativ Joshi, Sohini Chintala, Scott Niekum

首次发表
浏览论文内容

中文总结 AI 辅助

研究提出拉格朗日奖励增强框架,从含奖励与成本模型的约束目标出发,经对偶化将问题转化为一维凸问题,校准对偶变量获增强奖励,提升推理时对齐的有益性与无害性权衡。

中文摘要 AI 辅助

推理时对齐利用辅助奖励信号在解码时引导冻结语言模型,避免重复权重更新成本。现有方法通常优化单个标量分数,安全约束易被忽略或需手动调整惩罚。我们提出拉格朗日奖励增强(LARA),一种安全约束下的通用推理时对齐框架……

英文摘要

Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize a single scalar score, so explicit safety constraints must either be ignored or encoded through manually tuned penalties. We propose Lagrangian Reward Augmentation (LARA), a general inference-time alignment framework under safety constraints. Starting from a KL-regularized constrained objective with a reward model and a cost model, LARA dualizes the constraint and reduces the optimization problem to a one-dimensional convex problem over a nonnegative dual variable. Estimated on a small calibration set, this dual variable defines an augmented reward that can be used as a drop-in scoring signal within existing inference-time alignment methods. For sequence-level sampling methods, such as Best-of-N reranking, the calibrated dual variable corresponds to the solution of the expected-cost constrained problem. For token-level reward-guided decoding methods, the same construction yields a principled dual-calibrated heuristic rather than an exact constrained-policy guarantee. We evaluate LARA on both sequence-level and token-level inference-time alignment methods, and find that LARA improves the helpfulness-harmlessness tradeoff, with Best-of-N achieving the best performance among inference-time methods, approaching finetuning-based direct alignment baselines.

发表机构

  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

↑