发表机构
Beijing Jiaotong University; Weixin AI, Tencent Inc(北京交通大学; 腾讯微云人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多语言长视野工作流程中LLM智能体的表现,提出PolyWorkBench基准测试及结合多种评估方式的混合框架,发现最先进LLM智能体在多语言设置中性能降,强调联合建模语言变化和程序决策对智能体评估的重要性。
AI 中文摘要
大型语言模型(LLM)智能体在需要规划、工具使用和与外部环境交互的长视野任务中表现出强大性能。但大多数现有基准测试隐含假设单语言设置。现实应用常涉及统一工作流程中的多语言输入输出,多语言与智能体执行间的交互尚少被探索。本文引入PolyWorkBench,用于评估LLM智能体在多语言长视野工作流程中的表现。它包含五个领域的67个任务,提出结合结构分级、可执行验证和基于LLM的语义评估的混合框架。实验表明,与单语言对应物相比,最先进的LLM智能体在多语言工作流程设置中性能显著下降。分析表明多语言在推理和执行步骤中引入复合效应,凸显在智能体评估中联合建模语言变化和程序决策的重要性。
英文摘要
While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across extended trajectories. The interaction between multilinguality and long-horizon execution, however, remains underexplored. We introduce PolyWorkBench, a benchmark designed to evaluate LLM agents on multilingual, long-horizon workplace workflows. PolyWorkBench features 67 tasks across five core domains: commerce, knowledge work, legal analysis, localization, and manufacturing. Tasks are authored by the paper's authors from real-world data seeds and independently verified through a second-author audit. Agents must integrate heterogeneous multilingual inputs, execute iterative tool-use trajectories, and produce structured domain artifacts. To rigorously assess performance, we adopt Grade, a task-specific structural scoring rubric, as our primary ranking metric, and complement it with Pytest for executable state verification and LLM-as-Judge for semantic quality diagnostics. Benchmark evaluations reveal that agent performance varies substantially across languages and drops sharply on the harder cross-lingual tasks, and our analysis shows that multilingual execution exposes systematic failure modes across planning, tool interaction, and decision-making in long-horizon agents.
Comments17 Pages, 5 figures