发表机构
Black in AI Safety & Ethics (BASE)(黑人AI安全与伦理(BASE))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究审计多语言AI安全训练数据集的语言缺口,发现其与资源层级不完全相关,豪萨语翻译质量未达阈值,非洲语言缺乏特定危害类别覆盖,提出切片级审计方法及建议。
AI 中文摘要
大型语言模型提供商通常会引用覆盖十几种语言的多语言安全基准,作为其模型对非英语用户而言是安全的证据。我们表明,这些集合层面的覆盖范围主张在单个语言层面的检验中往往站不住脚。我们审计了21个资源,涵盖25个语言切片,其中根据我们的统计规则有20个可算作数据集,涉及三种分别代表低资源(豪萨语)、中等资源(斯瓦希里语)和高资源(法语)层级的语言,我们发现,在来源、注释可靠性、访问权限、危害分类覆盖范围以及数据复用方面存在缺口,这些缺口的模式部分但并非完全与资源层级相关。通过管道内的受控比较,我们发现豪萨语切片低于其论文自身的翻译质量接受阈值,而同一管道的斯瓦希里语输出则轻松达到该标准;这表明这些缺口是可测量和可解决的,而非固有的。我们进一步发现,在我们研究的两种非洲语言层级中,自我伤害和性内容类别均没有本族语覆盖,这是一种整体而非渐进的缺口,是纯资源层级解释无法预测的。我们将这些发现与多语言越狱鲁棒性中已记录的持续不对称性(单轮攻击大多已缓解,多轮攻击仍然有效)联系起来,认为这种不对称性在结构上与我们审计发现的训练和评估数据最薄弱的地方一致。我们贡献了一种可重复使用的切片级审计方法、跨层级实证比较,以及针对数据集创建者、模型提供商和 venues 的具体建议,旨在使“多语言覆盖”主张可验证而非仅为陈述。数据集:this https URL
英文摘要
Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users. We show that these collection-level coverage claims frequently do not survive inspection at the level of an individual language. Auditing 21 resources across 25 language slices, of which 20 count as datasets under our counting rules, spanning three languages chosen to represent low- (Hausa), mid- (Swahili), and high-resource (French) tiers, we find that gaps in provenance, annotation reliability, access, harm-taxonomy coverage, and data reuse recur in patterns that partially, but not fully, track resource level. Using a controlled within-pipeline comparison, we show a Hausa-language slice falling below its own paper's translation-quality acceptance threshold while the same pipeline's Swahili output clears the same bar comfortably; this is evidence that these gaps are measurable and addressable, not inherent. We further show that self-harm and sexual-content categories have no native-language coverage in either African-language tier we studied, a total rather than gradated gap that a purely resource-level account does not predict. We connect these findings to a documented, persistent asymmetry in multilingual jailbreak robustness (single-turn attacks largely mitigated, multi-turn attacks still effective), arguing that this asymmetry is structurally consistent with where our audit finds training and evaluation data thinnest. We contribute a reusable slice-level audit methodology, a cross-tier empirical comparison, and concrete recommendations for dataset creators, model providers, and venues aiming to make ``multilingual coverage'' claims verifiable rather than merely stated. Dataset: https://huggingface.co/datasets/ChialukaOnuoha/safety-slice-audit