WorldBench:面向多语言智能体的文化贴合基准
WorldBench: Culturally Grounded Benchmark for Multilingual Agents
- University of Edinburgh(爱丁堡大学)
- School of Informatics, University of Edinburgh(爱丁堡大学信息学院)
- School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出WorldBench多语言文化贴合基准,含1600个跨语言跨文化任务,引入CTS指标评估,发现前沿模型CTS仅49.2%,当前智能体在相关场景下仍较脆弱。
AI中文摘要:
尽管基于大语言模型(LLM)的智能体在复杂环境中解决多步骤任务的应用日益广泛,但现有基准很少测试状态保留、跨语言性能及现实落地场景的应用。为解决这些问题,我们提出WorldBench:一个全面的多语言基准,包含真实的、基于角色设定的日常工作流,智能体可通过结构化操作在沙箱中执行。WorldBench涵盖7种语言、8种文化下的1600个任务,经具备语言和文化专业知识的人工标注员反馈筛选优化。评估中,我们扩展了过往工作的指标,引入受限任务成功度(CTS),其结合自然语言指令与测试平台,通过确定性评估和大语言模型作为评判者的评估,对任务完成度、最小修改及其他补充指标打分。实验显示,前沿模型的CTS仅达49.2%,所有模型均在正确性与状态保留间存在巨大差距。由此表明,当前智能体在多语言智能体场景中仍较脆弱,尤其在长周期任务和状态保留约束下表现不佳。
英文摘要:
Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to realistic, grounded scenarios. To address these concerns, we present WorldBench: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions. WorldBench comprises 1,600 tasks across seven languages and eight cultures, filtered and refined through feedback from human annotators with language- and culture-specific expertise. For evaluation, we extend metrics from previous works and introduce Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and other complementary metrics through deterministic and LLM-as-a-Judge evaluations. Our experiments show that frontier models reach only 49.2% CTS, with all models demonstrating large gaps between correctness and environment preservation. We thereby show that current agents remain brittle in multilingual, agentic scenarios, especially for long-horizon tasks and under state-preservation constraints