arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PolyWorkBench:多语言长视野语言模型智能体基准测试

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows

Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou, Jinan Xu, Fandong Meng, Kaiyu Huang

arXiv 2607.06008首次发表:更新:

发表机构

Beijing Jiaotong University; Weixin AI, Tencent Inc(北京交通大学; 腾讯微云人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多语言长视野工作流程中LLM智能体的表现,提出PolyWorkBench基准测试及结合多种评估方式的混合框架,发现最先进LLM智能体在多语言设置中性能降,强调联合建模语言变化和程序决策对智能体评估的重要性。

AI 中文摘要

大型语言模型(LLM)智能体在需要规划、工具使用和与外部环境交互的长视野任务中表现出强大性能。但大多数现有基准测试隐含假设单语言设置。现实应用常涉及统一工作流程中的多语言输入输出,多语言与智能体执行间的交互尚少被探索。本文引入PolyWorkBench,用于评估LLM智能体在多语言长视野工作流程中的表现。它包含五个领域的67个任务,提出结合结构分级、可执行验证和基于LLM的语义评估的混合框架。实验表明,与单语言对应物相比,最先进的LLM智能体在多语言工作流程设置中性能显著下降。分析表明多语言在推理和执行步骤中引入复合效应,凸显在智能体评估中联合建模语言变化和程序决策的重要性。

英文摘要

While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across extended trajectories. The interaction between multilinguality and long-horizon execution, however, remains underexplored. We introduce PolyWorkBench, a benchmark designed to evaluate LLM agents on multilingual, long-horizon workplace workflows. PolyWorkBench features 67 tasks across five core domains: commerce, knowledge work, legal analysis, localization, and manufacturing. Tasks are authored by the paper's authors from real-world data seeds and independently verified through a second-author audit. Agents must integrate heterogeneous multilingual inputs, execute iterative tool-use trajectories, and produce structured domain artifacts. To rigorously assess performance, we adopt Grade, a task-specific structural scoring rubric, as our primary ranking metric, and complement it with Pytest for executable state verification and LLM-as-Judge for semantic quality diagnostics. Benchmark evaluations reveal that agent performance varies substantially across languages and drops sharply on the harder cross-lingual tasks, and our analysis shows that multilingual execution exposes systematic failure modes across planning, tool interaction, and decision-making in long-horizon agents.

Comments17 Pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑