发表机构
Huawei Technologies Co., Ltd; Shanghai Jiao Tong University; Northwestern University; Harbin Institute of Technology, Shenzhen; Shenzhen Loop Area Institute(华为技术有限公司; 上海交通大学; 西北大学; 哈尔滨工业大学(深圳); 深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM智能体数据生成领域提出两级框架与ACE视角,明确核心挑战是为演化的智能体与环境分配有效、非冗余的经验。
AI 中文摘要
大语言模型(LLM)智能体越来越依赖生成的交互数据来学习如何与外部环境交互。智能体数据生成必须在环境、任务、交互和成功信号之间保持一致性,同时产生有用而非仅仅数量庞大的经验。现有研究覆盖了多个智能体领域,但以领域为中心的组织方式和异构的评估方法往往掩盖了通用的生成机制,还将候选构建与验证、选择过程混为一谈。本研究为该领域开发了一个两级框架:首先,我们将智能体数据表示为一个通用的分解对象(E,q,τ,v),包含环境规范、任务信号、交互实现和可选验证器;我们根据生成范式的主要锚点和依赖结构对其进行组织。其次,我们通过准确率-复杂度-多样性(ACE)视角将生成过程形式化为约束分布设计:准确率确定接地且内部一致数据的可行支撑集,在该支撑集内,复杂度相对于声明的学习器和执行配置分配学习权重,而多样性控制数据的覆盖范围和冗余度。利用该框架,我们探究了现有工作如何验证生成的经验、构建和校准难度以及扩展行为覆盖范围。文献显示,研究正朝着基于执行的准确率、与学习器相关的复杂度以及超越表面变化或数据集规模的多样性方向转变。我们还通过ACE视角讨论了智能体数据生成的更广泛方向和新兴趋势,包括其对规模扩展、数据源、训练机制和自适应学习的影响。总体而言,核心挑战并非简单生成更多数据,而是随着智能体和环境的演变,持续分配有效、信息丰富且非冗余的经验。
英文摘要
LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while producing experience that is useful rather than merely abundant. Existing work spans many agent domains, but domain-centered organization and heterogeneous evaluation often obscure common generation mechanisms and conflate candidate construction with verification and selection. This work develops a two-level framework for the field. First, we represent agentic data as a common factorized object $(E,q,τ,v)$, comprising an environment specification, task signal, interaction realization, and optional verifier. We organize generation paradigms by their primary anchor and dependency structure. Second, we formulate generation as constrained distribution design through the Accuracy-Complexity-divErsity (ACE) lens. Accuracy establishes the feasible support of grounded and internally consistent data. Within this support, Complexity places learning mass relative to the capability of a declared learner and execution configuration, while divErsity controls coverage and redundancy of data. Using this framework, we explore how prior work verifies generated experience, constructs and calibrates difficulty, and expands behavioral coverage. The literature reveals a shift toward execution-grounded accuracy, learner-relative complexity, and diversity beyond surface variation or dataset size. We further discuss broader directions and emerging trends in agentic data generation through the ACE lens, including their implications for scaling, data sources, training regimes and adaptive learning. Overall, the central challenge is not simply to generate more data, but to continually allocate valid, informative, and non-redundant experience as agents and environments evolve.