AI 中文总结
针对扩散模型生成图像不可分层编辑、编码智能体视觉生成缺乏审美直觉的问题,本文提出可编辑视觉设计范式,结合VLM与图像生成模型,实现兼具精致审美与生产级可编辑性的视觉内容生成。
AI 中文摘要
尽管扩散基础模型如GPT-Image-2和Nano-Banana具有出色的视觉表现力,但其端到端生成的图像本质上是扁平化位图,文本易出错,无法进行分层后期编辑。相比之下,通过编码智能体(Coding Agent)实现的基于代码的视觉生成,虽能提供精确的布局控制和分层解耦能力,但仍存在全局审美直觉不足、编写复杂视觉资源难度大的问题。为解决上述问题,本文提出了一种由编码智能体驱动的新范式——可编辑视觉设计。我们将视觉语言模型(VLM)指定为创意大脑,负责需求理解、任务规划和审美判断,同时利用图像生成模型作为按需视觉世界模拟器,合成独立的视觉资源。该智能体在“先想象,后行动”的闭环工作流程下,生成独立资源、编写原生HTML/CSS代码,并根据视觉渲染反馈迭代优化设计。此外,智能体设计回放(Agent Design Replay)可忠实地复现类似专业人类设计师的创意与推理轨迹。最终,该系统生成具有解耦图层和真实文本的可编辑产物,使用户能在图形用户界面上进行直观的鼠标拖动和布局调整。在海报、信息图等场景的验证表明,该范式成功实现了精致的审美效果和生产级别的可编辑性。
英文摘要
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control and decoupled layers, yet remains constrained by a lack of global aesthetic intuition and the difficulty of coding complex visual assets. To address this, we propose Editable Visual Design, a new paradigm driven by a Coding Agent. We designate the VLM as the ``creative brain'' for requirement comprehension, task planning, and aesthetic judgment, while utilizing the image generation model as an on-demand ``visual world simulator'' to synthesize standalone visual assets. Operating under an ``imagine first, then act'' closed-loop workflow, the agent generates isolated assets, writes native HTML/CSS, and iteratively refines the design against visual rendering feedback. Furthermore, Agent Design Replay faithfully reproduces the creative and reasoning trajectory akin to that of professional human designers. Ultimately, the system delivers editable artifacts with decoupled layers and real text, enabling users to perform intuitive mouse dragging and layout adjustments on a graphical user interface. Validations on posters, infographics, and other scenarios show that this paradigm successfully achieves both refined aesthetics and production-grade editability.