arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体视觉生成:从生成模型到智能体控制

Agentic Visual Generation: From Generative Models to Agentic Control

Yinming Huang, Shuyuan Tu, Xi Yan, Jiahao Zhan, Zihan Yang, Zhen Xing, Hui Zhang, Tiehua Zhang, Yu-Gang Jiang, Zuxuan Wu

arXiv 2609.06758首次发表:更新:

发表机构

Fudan University; Shanghai Innovative Institute; CUHK; Alibaba Tongyi Lab; Tongji University(复旦大学; 上海创新研究院; 香港中文大学; 阿里通义实验室; 同济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出按控制器在生成过程中的直接控制范围划分L0-L4层级,以统一标准界定生成系统的智能体性,并据此分析图像、视频等多领域生成中控制器能力的演进。

AI 中文摘要

视觉生成正从通过单次调用使用的生成模型,演变为能够规划、选择工具、检查中间合成输出、修正失败并复用先前经验的智能体控制过程。在大多数现有系统中,控制器是LLM或VLM,而视觉生成模型充当工具或执行器。然而,现有工作缺乏一致的标准来确定生成系统何时变得具有智能体性。规划深度、工具使用、多角色协作和强化学习常被视为智能体性的证据,尽管它们都不必然决定控制器能够做出哪些生成决策。我们根据控制器在生成过程中能直接控制的内容来组织该领域。在L1条件控制中,控制器准备预定生成器的输入,但不控制执行哪种视觉操作。在L2执行控制中,它选择并调用实际的生成、编辑、渲染或其他内容修改操作。在L3结果自适应控制中,它观察中间结果,并利用该观察在当前任务内更改后续操作。在L4经验自适应控制中,它保留已完成任务的经验,并利用该经验更改未来任务的决策。L0固定支持单独表示生成器、编辑器、评估器、奖励模型、基准和固定流水线,这些没有部署做出生成级决策的控制器。这些层级描述的是逐步更广泛的决策范围,而非模型大小、系统复杂度、输出质量、工具或角色数量或训练方法。将该框架应用于图像、视频、编辑、3D、世界、幻灯片和用户界面生成,揭示了控制器能力如何演变以及其机制如何分布在各个层级。

英文摘要

Visual generation is evolving from generative models used through a single invocation into agentic control processes that can plan, select tools, inspect intermediate synthesized outputs, revise failures, and reuse prior experience. In most existing systems, the controller is an LLM or VLM, while visual generation models serve as tools or executors. However, existing work lacks a consistent criterion for determining when a generation system becomes agentic. Planning depth, tool use, multi-role collaboration, and reinforcement learning are often treated as evidence of agenticity, even though none of them necessarily determines which generation decisions the controller can make. We organize the field according to what the controller can directly control in the generation process. At L1 Conditioning Control, the controller prepares the input to a predetermined generator but does not control which visual operation is executed. At L2 Execution Control, it selects and invokes actual generation, editing, rendering, or other content-modifying operations. At L3 Outcome-Adaptive Control, it observes an intermediate outcome and uses that observation to change a subsequent operation within the current task. At L4 Experience-Adaptive Control, it retains experience from completed tasks and uses that experience to change decisions on future tasks. L0 Fixed Support separately denotes generators, editors, evaluators, reward models, benchmarks, and fixed pipelines without a deployed controller that makes generation-level decisions. These levels describe a progressively broader decision-making scope rather than model size, system complexity, output quality, tool or role count, or training method. Applying this framework across image, video, editing, 3D, world, slide, and user-interface generation reveals how controller capabilities have evolved and how their mechanisms are distributed across levels.

Commentsproject page: https://github.com/YinmingHuang/Awesome-agentic-visual-generation-model

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑