arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniaBench:跨多样场景对通用人工智能智能体进行基准测试

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

Chengyu Shen, Yujie Fu, Gangtao Xin, Yanheng Hou, Wenlong Fei, Guojie Zhu, Jiawei Li, Hongcheng Gao, Runming He, Zhen Hao Wong, Meiyi Qiang, Hao Liang, Zhao Cao, Hao Jiang, Chong Chen, Wentao Zhang

arXiv 2607.14989首次发表:更新:

AI 中文总结

研究针对大语言模型向通用智能体演变,现有基准测试有局限的问题,引入OmniaBench。通过多渠道获取场景知识形成分类法,构建可执行环境并合成任务,还引入能力分类法等。其数据集含多任务,对前沿模型构成挑战,能表征通用智能体能力边界。

AI 中文摘要

大语言模型正逐渐从文本生成器演变为能够理解用户请求、调用外部工具并通过交互完成复杂任务的通用智能体。然而,现有智能体基准测试往往局限于有限场景、工具生态系统或交互格式,难以系统地表征模型在异构应用设置中的能力。我们引入了OmniaBench,这是一个用于在具有明确状态空间的多样场景中评估通用智能体的基准测试。我们从应用商店、产品文档、行业资源、网络检索和人工完善中获取面向应用的场景知识,形成一个跨越消费、企业和教育领域的层次分类法,包含90个一级和354个二级领域。基于此分类法,我们构建可执行环境,并通过四种互补途径:有向无环图(DAG)、DAG-S、求解器和程序,合成单轮和多轮任务。OmniaBench还引入了一个十维能力分类法和八个组合原子难度因素,以支持细粒度评估和分析。生成的数据集包含1431个任务,以及一个具有挑战性的644个任务子集,旨在降低评估成本并减轻公开发布后全集的潜在污染。该基准测试对当前前沿模型提出了重大挑战,即使是Claude-Sonnet-5和GPT-5.6-Sol的总体通过率@1分数也仅分别为58.54和57.14。进一步分析揭示了各领域和能力之间的明显差异,以及在规划、约束维护和自适应校正方面的持续局限性。OmniaBench为表征通用智能体的能力边界提供了一个广泛且具有诊断性的基准测试。

英文摘要

Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tasks through interaction. However, existing agent benchmarks often focus on limited scenarios, tool ecosystems, or interaction formats, making it difficult to systematically characterize model capabilities across heterogeneous application settings. We introduce OmniaBench, a benchmark for evaluating general agents across diverse scenarios with explicit state spaces. We derive application-oriented scenario knowledge from app stores, product documents, industry resources, Web retrieval, and human refinement, forming a hierarchical taxonomy that spans ToC, ToB and ToE with 90 level-1 and 354 level-2 domains. Based on this taxonomy, we construct executable environments and synthesize single-turn and multi-turn tasks through four complementary routes: DAG, DAG-S, Solver, and Program. OmniaBench further introduces a ten-dimensional capability taxonomy and eight compositional atomic difficulty factors to support fine-grained evaluation and analysis. The resulting dataset contains 1,431 tasks, together with a challenging subset of 644 tasks designed to reduce evaluation cost and mitigate potential contamination of the full set after public release. The bench presents substantial challenges to current frontier models, with even Claude-Sonnet-5 and GPT-5.6-Sol achieving Overall Pass@1 scores of only 58.54 and 57.14, respectively. Further analyses reveal clear differences across domains and capabilities, as well as persistent limitations in planning, constraint maintenance, and adaptive correction. OmniaBench provides a broad and diagnostic benchmark for characterizing the capability boundaries of general agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑