arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35805cs.CLcs.AI

对齐预测:从训练数据预测错位

Alignment Forecasting: Predicting Misalignment From Training Data

发表机构纽约大学 · MATS · OpenAI
查看机构详情
  • NYU(纽约大学)
  • MATS
  • OpenAI

机构由 AI 辅助整理,请以论文原文为准。

Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky, Tomek Korbak

首次发表
浏览论文内容

中文总结 AI 辅助

提出对齐预测任务,在训练前预测微调数据引发的模型错位,并构建ALIGNMENTFORECASTBENCH基准,通过结合LLM评分和简单学习模型的脚手架,在SFT场景中实现高于随机水平的预测,并能过滤有害数据提升对齐效果。

中文摘要 AI 辅助

在具有细微缺陷的数据上训练语言模型有时会使模型广泛错位。仅凭表面检查数据往往无法确定这种错位是否会出现,而目前只有在训练之后通过审计结果模型才能发现。为了补充事后审计,我们引入了对齐预测(Alignment Forecasting)任务:在训练之前预测对齐失败。给定目标模型、微调数据集和失败模式(如欺骗或奉承),预测器输出微调会显著增加该失败模式的概率。为了衡量对齐预测的进展,我们引入了ALIGNMENTFORECASTBENCH基准,包含超过5000个预测问题,涵盖17个目标模型、32个数据集和16种失败模式。直接提示的前沿模型在ALIGNMENTFORECASTBENCH上表现不佳。因此,我们提出了一种预测脚手架,其中LLM读取数据集并评估其推动模型走向不当行为的强度和广度,然后一个简单的学习模型将该评分与失败模式的基础率以及目标模型的先前倾向相结合。这种预测显著高于随机水平,并且优于在任务上微调的模型以及一个允许查看较弱模型在同一数据上微调后行为的简单预测器。其信号还能标记出前沿模型分类器遗漏的有问题的训练示例。从真实的后训练数据(如UltraChat)中过滤掉这些示例,在我们的多项选择评估中大多数情况下会产生更对齐的模型,尽管在开放式对话中的益处尚不清楚。在预测能够可靠地指导实际训练数据整理之前,还需要更多进展,但我们的结果表明,在SFT设置中,预测许多训练前的对齐失败是可行的。

英文摘要

Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 datasets, and 16 failure modes. Frontier models prompted directly perform poorly on ALIGNMENTFORECASTBENCH. We therefore propose a forecasting scaffold in which an LLM reads the dataset and rates how strongly and broadly it pushes the model toward misbehavior, and a simple learned model combines that rating with the failure mode's base rate and the target model's prior tendency. This forecasts well above chance, and beats a model fine-tuned on the task and a simple forecaster allowed to see how weaker models behaved after fine-tuning on the same data. Its signals also flag problematic training examples that a frontier-model classifier misses. Filtering those examples out from real post-training data such as UltraChat results in more aligned models on our multiple-choice evaluation in most cases, though the benefit in open-ended conversations is unclear. More progress is needed before forecasts can reliably guide training data curation in practice, but our results suggest that forecasting many alignment failures before training can be tractable in the SFT setting.

↑