Arabic Safety Alignment as Selective Refusal: An Empirical Study of SFT, DPO, and Guard Calibration
阿拉伯语安全对齐作为选择性拒绝:SFT、DPO与防护校准的实证研究
机构 * American University of Beirut(贝鲁特美国大学)
专题命中 偏好对齐 :DPO(title,title_cn);alignment(title);safety(title);分类 cs.CL、cs.AI
AI总结 该研究针对阿拉伯语大语言模型的安全对齐问题,通过实证对比SFT、DPO等方法,提出需针对特定模型选择操作点以平衡有害提示拒绝与良性提示接受的权衡。
Comments The Fourth Arabic Natural Language Processing Conference(ArabicNLP 2026)