发表机构
Adobe Research; KAIST(奥多比研究院; 韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlowTool将基于工具的图像编辑视为流匹配问题,通过条件修正流直接建模工具参数分布,结合视觉-语言模型和扩散Transformer生成编辑计划,在多个基准上性能优异且推理延迟降低至少50倍。
AI 中文摘要
基于工具的图像编辑(图像修饰)通常使用自回归多模态大语言模型(MLLMs)来公式化,这些模型依次生成推理、工具选择和参数值。在这项工作中,我们提出了一种新颖的基于工具的图像编辑方法,将任务构建为流匹配问题。我们引入了FlowTool,一个框架,它直接建模以输入图像和用户指令为条件的高质量工具参数的分布,使用条件修正流。FlowTool结合了用于多模态理解的视觉-语言模型主干和一个扩散Transformer参数生成器,该生成器将高斯噪声转换为编辑计划。我们使用两阶段监督流匹配课程训练FlowTool,随后进行基于奖励的后训练。在MMArt-Bench、FlowTool-Eval、ArtEdit-Bench和MIT-Adobe5K上,FlowTool在基于参考的评估中显著优于专门的MLLM编辑代理和专有MLLM,同时在无参考评估中与专有模型保持竞争力。此外,FlowTool显著提高了推理效率,将延迟降低了至少50倍,同时所需内存比对比基线少近2倍。这些结果表明,基于工具的图像编辑可以有效地建模为对结构化连续编辑参数的条件生成,而无需自回归推理。
英文摘要
Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models (MLLMs) that sequentially generate reasoning, tool selections, and parameter values. In this work, we present a novel approach to tool-based image editing by framing the task as a flow matching problem. We introduce FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow. FlowTool combines a vision-language model backbone for multimodal understanding with a Diffusion Transformer parameter generator that transforms Gaussian noise into an editing plan. We train FlowTool with a two-stage supervised flow-matching curriculum, followed by reward-based post-training. Across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K, FlowTool achieves significantly stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, while remaining competitive with proprietary models under reference-free evaluation. Moreover, FlowTool significantly improves inference efficiency, reducing latency by at least $50\times$ while requiring nearly $2\times$ less memory than the compared baselines. These results demonstrate that tool-based image editing can be effectively modeled as conditional generation over structured continuous editing parameters, without autoregressive reasoning.