发表机构
American University of Beirut(贝鲁特美国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对阿拉伯语大语言模型的安全对齐问题,通过实证对比SFT、DPO等方法,提出需针对特定模型选择操作点以平衡有害提示拒绝与良性提示接受的权衡。
AI 中文摘要
阿拉伯语大语言模型必须拒绝有害提示,同时避免过度拒绝良性或敏感提示,但单一拒绝率掩盖了这一权衡。我们采用良性拒绝率B和有害提示拒绝率H进行评估,其中H衡量拒绝而非有害合规性。在5种具备阿拉伯语能力的模型及针对完整人工编写的AraSafe数据集的130次运行中,仅拒绝式监督微调(SFT)会退化为全面拒绝,而选定的混合SFT配置在B为14%至23%时,H达到90%至93%;4种选定配置在全部3次运行中均超过H=90%的目标,而Fanar模型在3次运行中的2次达到该目标。直接偏好优化(DPO)与推理防护会在不同模型中以不同方式改变B和H,而非作为统一升级。在一项盲法300次响应审核中,标注员二元拒绝一致性为89.0%(kappa=0.78);Qwen3Guard和Aya Expanse 32B的准确率分别达到88.7%和91.0%,无确凿配对差异。选定的SFT提升了所有5种模型对阿拉伯语(Arabizi)的H值,但无一种达到90%,显示现代标准阿拉伯语的迁移仅为部分。总体而言,结果支持针对特定模型的操作点选择:设定部署目标并仅保留能提升该目标的干预措施。
英文摘要
Arabic large language models must refuse harmful prompts without over-refusing benign or sensitive prompts, yet a single refusal rate hides this trade-off. We evaluate it using benign refusal B and harmful-prompt refusal H, where H measures refusal rather than harmful compliance. Across five Arabic-capable models and 130 runs on the full human-written AraSafe set, refusal-only supervised fine-tuning (SFT) collapses toward blanket refusal, whereas selected mixed-SFT configurations reach H = 90% to 93% at B = 14% to 23%; four selected configurations exceed the H = 90% target in all three runs, while Fanar does so in two of three. Direct Preference Optimization (DPO) and inference guards change B and H differently across models rather than acting as uniform upgrades. In a blinded 300-response audit, annotator binary-refusal agreement is 89.0% (kappa = 0.78); Qwen3Guard and Aya Expanse 32B reach 88.7% and 91.0% accuracy, respectively, with no conclusive paired difference. Selected SFT raises H on Arabizi for all five models, but none reaches 90%, showing only partial transfer from Modern Standard Arabic. Overall, the results support model-specific operating-point selection: set a deployment target and retain only interventions that improve it.
CommentsThe Fourth Arabic Natural Language Processing Conference(ArabicNLP 2026)