发表机构
Meta Superintelligence Labs(元超级智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人员推出OmnilingualGAIA2基准,发现前沿AI智能体存在8.8-18.4分的跨语言差距,提出多语言评估应成为全球部署智能体的标准评估环节。
AI 中文摘要
智能体基准旨在衡量AI智能体在现实多工具环境中的规划、搜索、执行与恢复能力,但几乎全部以英语开展。随着AI智能体面向全球多语言用户部署,英语下测得的智能体能力能否迁移至其他语言仍是悬而未决的问题。我们推出OmnilingualGAIA2,它是GAIA2智能体基准的机器翻译扩展版(含部分人类专家验证),覆盖五种书写系统的十种目标语言,并搭配本地化且经人工校准的多语言验证器。通过评估七种前沿开源权重智能体,我们发现存在8.8-18.4个pass@3分的通用跨语言差距,该差距在幅度上具有智能体不对称性,集中于工具编排而非定量推理,且不会随模型规模缩小。分层误差归因分析将该差距分解为主要由模型驱动(占55%),而受翻译污染的下限仅占场景-语言对的6.4%。人类专家语言分析进一步指出,形态线索丢失与歧义放大是非拉丁文字语言中的主要失效机制。我们的结果表明,多语言智能体评估必须成为面向全球部署智能体报告协议的标准组成部分。
英文摘要
Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question. We introduce OmnilingualGAIA2, a machine-translated expansion (with partial human- expert validation) of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier. Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale. A stratified error attribution decomposes the gap as predominantly model-driven (55%), with a bounded translation-contamination floor of only 6.4% of scenario-language pairs. Human-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non-Latin-script languages. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.