arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22226cs.CLcs.LGstat.ML

Swiss-Knife:一种解码时重配置外部化多目标对齐框架

Swiss-Knife: A Framework for Reconfigurable Externalised Multi-Objective Alignment at Decode Time

  • BITS Pilani Hyderabad Campus(比拉理工学院海得拉巴校区)
  • RV College of Engineering(RV工程学院)
  • Google(谷歌)
  • Amazon(亚马逊)
  • Apple(苹果)
  • Canva
  • Pragya Lab, BITS Pilani Goa Campus(Pragya实验室,比拉理工学院果阿校区)

机构由 AI 辅助整理,请以论文原文为准。

Agnibh Karmakar, Mayur Parvatikar, Shreyash Dhoot, Amit Dhanda, Aman Chadha, Kapil Wanaskar, Vinija Jain, Amitava Das

AI总结:

Swiss-Knife提出一种解码时外部化多目标对齐框架,通过可热插拔刀片、批量归一化和成对聚合实现可重配置对齐,在有用性/诚实性/无害性上达到最佳平衡,且无需梯度计算。

AI中文摘要:

解码时对齐方法通过使用外部奖励对候选续写进行评分并选择最大化者,来引导冻结的语言模型。我们认为这种共享设计是更大空间中的一个退化点。我们引入了Swiss-Knife,一个用于外部化多目标对齐的框架,其中对齐规范是一等运行时对象:可热插拔的评分刀片、批量归一化器、成对聚合算子和选择规则。六个公理刻画了可接受的聚合算子,我们证明了一个表示定理:满足这些公理的每个算子都具有形式$R_i = \sum_{j \neq i} g((\mu_i - \mu_j)/s(\sigma_i,\sigma_j))$,这是一个包含probit和logistic比较规则以及逐点argmax作为命名坐标的双参数族。在该族内,成对聚合在对抗性奖励污染下是Lipschitz稳定的,而argmax则不是,并且候选批量归一化(CBN)使权重单纯形在奖励模型仅被识别的重缩放下保持不变。我们的参考实例将DPO-LoRA刀片与不确定性感知的成对锦标赛配对。扫描有用性/诚实性/无害性单纯形,它在六种解码时方法中达到了最佳平衡前沿(调和F1为0.797,而最强基线为0.750,p < 10^-14),具有最低的拒绝率和最高的有用性,并且在没有梯度计算的情况下在0.05毫秒内重新配置其目标。消融CBN损失0.217 F1,元数据证实了预测的机制:最低方差刀片保留了其名义33%影响力的9%,崩溃遵循预测的顺序。

英文摘要:

Decode-time alignment methods steer a frozen language model by scoring candidate continuations with an external reward and selecting the maximiser. We argue that this shared design is a single degenerate point in a much larger space. We introduce Swiss-Knife, a framework for externalised multi-objective alignment in which the alignment specification is a first-class runtime object: hot-swappable scoring blades, a batch normaliser, a pairwise aggregation operator, and a selection rule. Six axioms characterise the admissible aggregation operators, and we prove a representation theorem: every operator satisfying them has the form $R_i = \sum_{j \neq i} g((μ_i - μ_j)/s(σ_i,σ_j))$, a two-parameter family containing probit and logistic comparison rules and pointwise argmax as named coordinates. Within it, pairwise aggregation is Lipschitz-stable under adversarial reward contamination while argmax is not, and Candidate-Batch Normalization (CBN) makes the weight simplex invariant to the rescalings under which reward models are only ever identified. Our reference instantiation pairs DPO-LoRA blades with an uncertainty-aware pairwise tournament. Sweeping the helpfulness/honesty/harmlessness simplex, it attains the best balanced frontier of six decode-time methods (harmonic $F_1$ 0.797 vs. 0.750 for the strongest baseline, $p < 10^{-14}$) with the lowest refusal rate and highest helpfulness of any arm, and reconfigures its objectives in 0.05 ms with no gradient computation. Ablating CBN costs 0.217 $F_1$, and the metadata confirms the predicted mechanism: the lowest-variance blade retains 9% of its nominal 33% influence, and the collapse follows the predicted ordering.

补充信息

↑