arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

技能问题:大型语言模型的技能是否具有语言不变性?

Skill Issue: Are Skills Language-Invariant in LLMs?

Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen

arXiv 2608.25832首次发表:更新:

发表机构

A*STAR; Weizmann Institute of Science; MIT-IBM Watson AI Lab; University of Cambridge; EleutherAI(新加坡科技研究局; 魏茨曼科学研究所; 麻省理工学院-IBM沃森人工智能实验室; 剑桥大学; EleutherAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究以多语言自博弈方法量化大型语言模型的跨语言技能不一致性,发现同一模型在不同语言下博弈实力差异显著,语言可影响决策阶段,技能差异是多语言模型开发的主要障碍。

AI 中文摘要

大型语言模型跨语言访问知识的表现不一致,但它们在与不同语言交互时的技能集差异程度如何?本研究从知识和通用基准性能之外的正交维度量化跨语言技能不一致性,方法是通过多语言自博弈:同一模型的两个实例在基于文本的游戏中竞争,每个实例通过不同语言的接口交互。由于模型、对手、规则、状态空间和可用动作保持固定,该设置可分离语言对模型实际行为的影响。我们构建了TextArena的多语言扩展版本,在涵盖空间推理、不完全信息、资源分配和重复交互的6款游戏、8种语言上评估了3个开放权重模型。研究发现,同一模型在不同语言下可表现出显著不同的博弈实力,胜负差、无效动作和策略倾向存在系统性差异。详细分析揭示了空间推理、基于卡牌的决策和最优走法选择中存在的语言特定失效情况。在部分场景中,仅改变中间推理语言即可恢复大部分损失的性能,表明语言可影响决策过程的不同阶段。这些结果表明,技能差异是开发真正多语言模型的可测量主要障碍,更好地理解这些差异可帮助我们设计出在各语言上表现更均衡的模型。

英文摘要

Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance. We do this via multilingual self-play: two instances of the same model compete in a text-based game, each interacting through a different language interface. Since the model, opponent, rules, state space, and available actions remain fixed, this setting isolates the effect of language on the model's realized behavior. We build a multilingual extension to TextArena and evaluate three open-weight models across eight languages and six games covering spatial reasoning, imperfect information, resource allocation, and repeated interaction. We find that the same model can exhibit markedly different playing strength across languages, with systematic variation in win--loss margins, invalid actions, and strategic tendencies. Detailed analyses reveal language-specific failures in spatial reasoning, card-conditioned decisions, and optimal move selection. In some settings, changing only the intermediate reasoning language recovers much of the lost performance, suggesting that language can affect different stages of the decision process. These results show that skill discrepancies are a measurable major roadblock in the development of truly multilingual models. Better understanding these discrepancies can help us design models that perform more equitably across languages.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑