arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从规范框架到对齐数据:构建与评估SFT和偏好数据

From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data

Husrev Taha Sencar, Rezart Beka, Danish Naeem, Seda Ozalkan, Majd Hawasly, Ji Lucas, Ala AlFuqaha, Mohamed Abdallah, Recep Senturk

arXiv 2609.35201首次发表:更新:

发表机构

Qatar Computing Research Institute, HBKU; College of Islamic Studies, HBKU; Ibn Haldun University; College of Science and Engineering, HBKU(卡塔尔计算研究所,哈马德·本·哈利法大学; 伊斯兰研究学院,哈马德·本·哈利法大学; 伊本·哈尔顿大学; 科学与工程学院,哈马德·本·哈利法大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种专家驱动方法,将伊斯兰伦理规范转化为SFT和偏好数据,通过受控实验验证其提升模型对齐效果,且不损害通用能力。

AI 中文摘要

将语言模型与特定的规范框架对齐,需要将抽象原则转化为具体的示例和偏好信号,模型才能从中学习。我们提出了一种由专家驱动的方法来构建此类对齐数据,并将其应用于一个植根于伊斯兰伦理、神学和法理学传统的规范框架。在大约一年的时间里,七位领域专家系统地探测了语言模型以识别对齐缺陷,整理了期望的响应,并从模型输出和专家判断中构建了偏好对。由此产生的阿拉伯语-英语数据集包含约2.8K个监督微调(SFT)示例和5.4K个偏好对,涵盖了广泛的规范领域。我们通过受控的后训练实验评估了这些数据集,比较了基线模型与仅包含整理的SFT数据以及同时包含SFT和偏好数据的模型。在150个单独构建的提示的盲法专家评估中,使用整理的SFT数据训练的模型在51.3%的评估者判断中优于基线模型,而相反方向的比例为14.4%(提示级别p<.001)。添加偏好数据导致差异较小,使用两个数据集训练的模型在28.0%的判断中优于仅使用SFT的模型,而相反方向的比例为20.9%;该差异在提示级别不具有统计学显著性(p=.166)。标准阿拉伯语和英语基准测试显示,通用能力没有广泛退化。这些结果证明了如何将专家定义的规范原则系统地操作化为对齐数据,并通过受控的模型训练进行评估。

英文摘要

Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions. Over approximately one year, seven domain experts systematically probed language models to identify alignment deficiencies, curated desired responses, and constructed preference pairs from model outputs and expert judgments. The resulting Arabic-English datasets contain approximately 2.8K supervised fine-tuning (SFT) examples and 5.4K preference pairs spanning a broad range of normative domains. We evaluate the datasets through controlled post-training experiments comparing a Baseline model with models incorporating the curated SFT data alone and both the SFT and preference data. In blind expert evaluation on 150 separately constructed prompts, the model trained with the curated SFT data was preferred over the Baseline in 51.3% of assessor judgments, compared with 14.4% in the opposite direction (p < .001 at the prompt level). Adding the preference data resulted in a smaller difference, with the model trained with both datasets preferred over the SFT model in 28.0% of judgments versus 20.9% in the opposite direction; this difference was not statistically significant at the prompt level (p = .166). Standard Arabic and English benchmarks show no broad degradation in general-purpose capabilities. These results demonstrate how expert-defined normative principles can be systematically operationalized into alignment data and evaluated through controlled model training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑