AI 中文总结
研究大型语言模型智能体全场景扩展问题,提出AgentOmnia框架,结合多种技术构建环境、工具和任务,通过多种方式支持训练后处理,能将多个基准测试通过率和宏平均率大幅提高,实现广泛改进并为自我进化提供证据。
AI 中文摘要
大型语言模型智能体发展迅速,但在跨领域、能力、任务难度和交互设置方面进展仍不连贯。我们将此视为全场景智能体扩展问题,并提出了AgentOmnia框架,用于协调面向消费者(ToC)、面向企业(ToB)和面向员工(ToE)应用中的任务空间定义、数据合成、训练后处理、评估和改进。一个可扩展的领域x能力x原子难度分类法对这些阶段进行了对齐,并通过OmniaBench实现细粒度诊断。AgentOmnia将双向环境-任务合成与工具依赖、程序结构和基于求解器的管道相结合,构建了5018个有状态环境、255375个工具和52361个任务。程序、求解器和验证器提供正确性信号,同时监督微调、在线智能体强化学习和回滚课程支持训练后处理。评估失败会转化为针对性自我进化的产品需求文档(PRD)。从Qwen3-30B-A3B-Thinking-2507开始,AgentOmnia将OmniaBench具有挑战性子集的通过率从9.16%提高到37.11%,并将OmniaBench、$\tau^2$-Bench、DeepPlanning和VitaBench的宏平均率从22.86%提高到41.69%。在统一协议下,它在OmniaBench上领先于经过评估的智能体训练后基线,并保持最高的四基准宏平均率。它在所有四个基准上也超过了Qwen3-235B-A22B-Thinking-2507,在宏平均上超过了Qwen3.5-35B-A3B。收益涵盖三个应用分类、十个能力维度、八个原子难度因素以及90个一级领域中的76个,表明是广泛而非特定类别的改进。一项单轮研究为PRD引导的自我进化提供了初步证据,推动在更大规模和工业环境中进行验证。
英文摘要
Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-scenario agentic scaling and present AgentOmnia, a framework coordinating task-space definition, data synthesis, post-training, evaluation, and improvement across To-Consumer (ToC), To-Business (ToB), and To-Employee (ToE) applications. An extensible Domain x Capability x Atomic Difficulty taxonomy aligns these stages and enables fine-grained diagnosis with OmniaBench. AgentOmnia combines bidirectional environment-task synthesis with tool-dependency, program-structured, and solver-based pipelines, constructing 5,018 stateful environments with 255,375 tools and 52,361 tasks. Programs, solvers, and verifiers provide correctness signals, while supervised fine-tuning, online agentic reinforcement learning, and a rollback curriculum support post-training. Evaluation failures translate into Product Requirement Documents (PRDs) for targeted self-evolution. Starting from Qwen3-30B-A3B-Thinking-2507, AgentOmnia raises the pass rate on the OmniaBench challenging subset from 9.16% to 37.11% and the macro-average across OmniaBench, $τ^2$-Bench, DeepPlanning, and VitaBench from 22.86% to 41.69%. Under a unified protocol,it leads the evaluated agentic post-trained baselines on OmniaBench and retains the highest four-benchmark macro-average. It also surpasses Qwen3-235B-A22B-Thinking-2507 on all four benchmarks and exceeds Qwen3.5-35B-A3B on the macro-average. Gains span three application splits, ten capability dimensions, eight atomic-difficulty factors, and 76 of 90 level-1 domains, indicating broad rather than category-specific improvement. A one-round study provides initial evidence for PRD-guided self-evolution, motivating validation at larger scales and in industrial settings.
Comments69 pages, 18 figures, 13 tables