arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于价值表征的对齐泛化预测

Predicting Alignment Generalization with Value Representations

Andy Liu, Mehar Bhatia, Karolina Stanczak, Mona Diab, Vered Shwartz, Daniel Fried

arXiv 2610.12410首次发表:更新:

发表机构

Carnegie Mellon University; Mila - Quebec AI Institute; McGill University; ETH Zurich; ETH AI Center; University of British Columbia; Vector Institute(卡内基梅隆大学; 米拉-魁北克人工智能研究所; 麦吉尔大学; 苏黎世联邦理工学院; 苏黎世联邦理工学院人工智能中心; 不列颠哥伦比亚大学; 矢量研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出对齐泛化预测任务,发现基于模型情境价值应用激活的表征优于文本描述表征,还构建了首个基于经验泛化动态的LLM价值分类体系。

AI 中文摘要

大语言模型(LLM)开发者会对模型进行后训练,使其展现出亲社会价值与行为特质,这些特质被列举在对齐目标中。然而,尽管近期的后训练进展已使模型在对齐评估中获得高分,但在狭窄行为集合上训练模型仍会以意外方式影响其在未见情境与环境中的行为。本文中,我们确立了对齐泛化预测任务,即预测针对给定价值微调模型如何改变其在大量保留价值上的行为。我们对现代对齐目标中包含的66种价值的对齐泛化效应开展大规模分析,并在对齐泛化预测任务上对表征技术进行基准测试。我们发现,基于模型在情境中应用价值时的激活表征,显著优于基于价值文本描述的方法。具体而言,最佳的激活基方法与我们的泛化矩阵达到0.45的相关性,而描述基基线仅为0.05。随后,我们展示了对齐泛化预测表征在下游任务中的适用性:利用这些表征衡量多价值对齐目标中各价值的相似性,我们发现该相似性与模型鲁棒性显著相关。最后,我们展示了指向共享、模型独立价值空间的初步证据,并用其开发首个基于经验泛化动态的LLM价值分类体系。本研究证明了研究LLM中价值泛化的重要性,及其在更具实证性的模型行为设计与训练中的应用价值。

英文摘要

LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values. We conduct a large-scale analysis of alignment generalization effects across 66 values found in modern alignment targets, and benchmark representational techniques on the alignment generalization prediction task. We find that representations based on model activations when applying values in context significantly outperform methods based on textual descriptions of the values. Specifically, the best activations-based methods achieve correlations of 0.45 with our generalization matrix, compared with 0.05 from description-based baselines. We then show the applicability of representations that predict alignment generalization toward downstream tasks by using them to measure how similar the values in a multi-value alignment target are, which we find is significantly correlated with model robustness. Finally, we show initial evidence towards a shared, model-independent value space, which we use to develop the first taxonomy of LLM values grounded in empirical generalization dynamics. Our work demonstrates the importance of studying value generalization in LLMs and its application toward the more empirical design and training of model behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑