发表机构
Alibaba Token Hub, Alibaba Group; University of Illinois Urbana-Champaign; Northeastern University; Monash University; Ant International(阿里巴巴集团阿里通义实验室; 伊利诺伊大学厄巴纳-香槟分校; 东北大学; 莫纳什大学; 蚂蚁国际)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出BabelFlow工作流构建多语言基准BabelArena,覆盖23种语言,揭示低资源语言智能体在工具使用、控制流和token效率上的显著差距,为多语言智能体研究奠定基础。
AI 中文摘要
大型语言模型(LLM)智能体日益通过工具使用以及与用户和环境的交互来执行多步骤工作流。然而,当前的智能体评估主要以英语为中心,限制了我们对智能体在多语言环境中能力的理解。我们引入了BabelFlow,一个基准通用的智能体工作流,它通过分析运行时依赖关系、协调保持结构不变的翻译,并将多层验证与人工审查相结合,将现有智能体基准适配到新语言,从而保留任务和评估语义。利用BabelFlow,我们构建了BabelArena,一个任务对齐的基准,包含来自四个基准家族、13个领域和23种语言的702个规范任务的16,146个实例。使用五个前沿模型的实验表明,没有一个模型在所有基准家族中占据主导地位,且跨语言差异远超任务成功率的范畴。低资源语言表现出不同的失败模式,其中工具使用和控制流错误的比例更大,而不仅仅是答案质量错误,这表明在这些语言的资源水平上,可靠任务执行存在差距。在相同任务上,低资源语言中的智能体消耗的token数也远多于英语(输入量最多约为英语的两倍),而交互长度并未成比例增加,并且在需要结构化输出的任务上,语言一致性进一步下降,其中切换 overwhelmingly 指向英语。我们相信BabelArena为推进可靠且高效的多语言智能体研究奠定了基础。
英文摘要
Large language model (LLM) agents increasingly execute multi-step workflows through tool use and interaction with users and environments. However, current agent evaluations are largely English-centric, limiting our understanding of agent capabilities in multilingual settings. We introduce BabelFlow, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-preserving translation, and combining multi-layer verification with human review to preserve task and evaluation semantics. Using BabelFlow, we construct BabelArena, a task-aligned benchmark comprising 16,146 instances derived from 702 canonical tasks across four benchmark families, 13 domains, and 23 languages. Experiments with five frontier models show that no single model dominates across benchmark families and that cross-language disparities extend well beyond task success. Lower-resource languages exhibit distinct failure patterns, with larger shares of tool-use and control-flow errors rather than answer-quality errors alone, pointing to gaps in reliable task execution across the resource levels of these languages. On the same tasks, agents in low-resource languages also consume substantially more tokens than in English (up to roughly twice the input) without proportional increases in interaction length, and language consistency degrades further on tasks requiring structured output, where switches are directed overwhelmingly toward English. We believe BabelArena provides a foundation for advancing research on reliable and efficient multilingual agents.
Comments20 pages, 11 tables, and 7 figures