困难负样本揭示容易负样本所掩盖的问题:跨语言有害性表征在困难负样本下随资源层级退化
Hard Negatives Reveal What Easy Negatives Hide: Cross-Lingual Harmfulness Representations Degrade with Resource Tier Under Hard Negatives
浏览论文内容
中文总结 AI 辅助
本研究证明跨语言有害性表征迁移依赖于负样本类型:在困难负样本下,低资源语言的有害性表征显著退化,表明仅用容易负样本评估不能证明表征跨语言存续。
中文摘要 AI 辅助
大语言模型的安全对齐主要使用英语进行训练,近期研究报道其底层的有害性表征能够跨越语言迁移:英语训练的探针在低资源语言中区分有害提示与无害提示的能力几乎与英语相当。这一结果被视作跨语言拒答失败主要反映校准问题而非表征质量的证据。我们证明这一结论取决于负样本的选择。在覆盖三个资源层级的九种语言中,当无害提示来自无关分布(容易负样本)时,我们复现了近完美的迁移(AUROC > 0.98)。当使用XSTest对比提示(良性但表面与有害请求相似,即困难负样本)时,迁移在低资源语言中崩溃,而在高资源语言中基本保持稳定。在Qwen2.5-7B-Instruct上,平均AUROC下降从英语的0.003增加到高资源语言的0.017、中资源语言的0.042和低资源语言的0.276。该模式在Aya Expanse上得到复现。回译chrF对照以及在三种语言上的匹配chrF比较降低了翻译质量解释该效应的可能性。在控制chrF后崩溃仍然存在(偏相关r = 0.70,p = 0.03)。分词器fertility与崩溃相关,并解释了资源层级效应的一部分,但并非全部。结果表明,容易负样本的迁移可以与困难负样本下的显著退化共存。因此,仅靠容易负样本评估无法确立有害性表征在翻译后仍然存续。
英文摘要
Safety alignment in large language models is trained primarily in English, and recent work reports that the underlying harmfulness representation survives translation: English-trained probes separate harmful from harmless prompts almost as well in low-resource languages as in English. This has been taken as evidence that cross-lingual refusal failures mainly reflect calibration rather than representation quality. We show that this conclusion depends on the choice of negative examples. Across nine languages spanning three resource tiers, we replicate near-perfect transfer (AUROC > 0.98) when harmless prompts come from an unrelated distribution (easy negatives). With XSTest contrast prompts, which are benign but surface-similar to harmful requests (hard negatives), transfer collapses in low-resource languages while remaining largely stable in high-resource languages. On Qwen2.5-7B-Instruct, mean AUROC drop increases from 0.003 in English to 0.017 in high-resource, 0.042 in mid-resource, and 0.276 in low-resource languages. The pattern replicates on Aya Expanse. Back-translation chrF controls and a matched-chrF comparison across three languages reduce the likelihood that translation quality explains the effect. The collapse remains after controlling for chrF (partial r = 0.70, p = 0.03). Tokenizer fertility correlates with the collapse and explains part of the resource-tier effect, but not all of it. The results show that easy-negative transfer can coexist with substantial degradation under hard negatives. Easy-negative evaluation alone therefore cannot establish that the harmfulness representation survives translation.
发表机构
- Birla Institute of Technology and Science, Pilani, Hyderabad Campus(比拉理工学院皮拉尼分校海得拉巴校区)
机构由 AI 辅助整理,请以论文原文为准。