公平事实核查:用RoSh缩小LLM事实判断中的跨语言差距
Fair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoSh
- Shifa Tameer-e-Millat University(希法塔米尔米拉特大学)
- BRAINS, Brandenburg Research Center for Applied Intelligent Systems(勃兰登堡应用智能系统研究中心)
- Max Planck Institute for Security and Privacy(马克斯·普朗克安全与隐私研究所)
- University of Cambridge(剑桥大学)
- The Pennsylvania State University(宾夕法尼亚州立大学)
- GISMA University of Applied Sciences(GISMA应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LLM事实判断的跨语言差距问题,提出RoSh方法,通过无训练的残差流平移旋转,平均缩小75%的差距,显著提升低资源语言表现。
AI中文摘要:
社交媒体上的错误信息仍然是一个严重问题,越来越多的人通过询问语言模型而非事实核查员来解决这一问题。模型是否能可靠地判断此类主张尚存争议;而模型是否能以人们提问所用的每种语言同样出色地判断这些主张,则几乎无人问津。我们测试了来自五个家族的八个模型,规模从3B到70B不等,针对1500条以八种语言以相同形式存在的百科全书式事实主张。在每个模型上,英语的判断效果都优于其他所有语言,且差距在最小模型上最为显著,其中Llama-3B在阿拉伯语上的表现不比随机猜测好。现有的补救方法要么在更多多语言数据上重新训练,要么拟合语言表示之间的无约束映射,而两者都没有追问模型是否已经掌握答案而只是未能表达出来。在很大程度上确实如此:线性探针能从模型未能表达的激活中恢复出真相。我们提出RoSh,一种对残差流进行逐语言平移和旋转的方法,在三个层上以闭式形式计算,无需训练,也不修改任何权重。它改善了每个模型,平均缩小了75%的差距,在模型表现最差的地方帮助最大:Llama-3B在阿拉伯语上从随机水平提升到接近英语水平,且用英语正确回答的主张中,有五分之一的损失在翻译中得以挽回。剩下的问题不再是读出失败:之后,头部从英语之外编码的信息中恢复的内容与从英语中恢复的一样多。在同一对数据上拟合的无约束映射低于未触及的基线,因此正交约束起了作用,且每个模型都通过了打乱对应关系的对照以及另外十项对照。在与最接近的推理时方法——潜空间干预——使用其自身数据和指标代码运行的两个基准上,RoSh的增益是其五到十三倍。
英文摘要:
Misinformation on social media remains a critical problem, and more and more people settle it by asking a language model instead of a fact checker. Whether models judge such claims reliably is debated; whether they judge them equally well in every language people ask in has gone almost unasked. We test eight models from five families, 3B to 70B, on 1,500 encyclopedic factual claims that exist in identical form in eight languages. English is judged better than every other language on every model, and the gap is widest on the smallest ones, where Llama-3B on Arabic is no better than guessing. Existing remedies retrain on more multilingual data or fit an unconstrained map between language representations, and neither asks whether the model already holds the answer and simply fails to say it. It largely does: a linear probe recovers the truth from the very activations the model fails to express. We propose RoSh, a per-language shift and rotation of the residual stream, computed in closed form at three layers, with no training and no weight modified. It improves every model and closes 75% of the gap on average, helping most where the model was worst: Arabic on Llama-3B goes from chance to nearly the English level, and a fifth fewer of the claims answered correctly in English are lost in translation. What remains is no longer a read-out failure: afterwards the head recovers as much of what is encoded outside English as it does in English. An unconstrained map fitted on the same pairs falls below the untouched baseline, so the orthogonality constraint is doing the work, and every model clears a scrambled-correspondence control and ten further controls. On the two benchmarks of the closest inference-time method, latent-space intervention, run with its own data and metric code, RoSh's gains are five to thirteen times larger.