发表机构
University of Surrey(萨里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ProgressNet提出无需训练的推理时机制,使冻结的文本到图像模型能渐进式跟随草图绘制与提示修改,每轮约一秒,在FS-COCO上FID几乎不变,优于现有方法。
AI 中文摘要
人类是渐进式绘图的:先画几笔,看一眼结果,擦除一笔,修改提示词。图像生成器并非如此工作。它们通常接受一张完成的草图,并在单次生成中产生图像,因此每次编辑都会重新开始生成图片,而那些能在多轮中保持状态的模型是由文本驱动的,无法接受笔画输入,并且速度太慢,无法用于绘图。我们提出了ProgressNet,一个无需训练的框架,它能让冻结的文本到图像模型跟随逐步展开的绘图会话:笔画被添加和擦除,提示词被修改,图像以每轮约一秒的速度保持同步。它不需要新参数,因为冻结模型已经具备了渐进式生成器所需的一切:一条能记住上一轮状态的路径,能在不冻结结构的情况下传递外观的层,以及一个指示对未完成草图信任程度的内部信号;三种推理时机制(先前概念记忆、层选择性K/V注入和带状自适应控制)依次利用这些能力。随着草图逐渐填充,所有现有方法都会退化,在FS-COCO上,FLUX+ControlNet基线的FID在完成度从10%到100%之间翻倍,而ProgressNet的FID几乎不变;它在三个草图领域保持了强大的保真度和渐进连贯性,并且用户对它的偏好超过了五个竞争对手,在擦除场景中优势最为明显。
英文摘要
Humans draw progressively: a few strokes, a look at the result, a stroke erased, a prompt revised. Image generators do not work this way. They typically take a finished sketch and produce the image in a single pass, so every edit starts the picture again, and the models that do keep state across turns are driven by text, cannot take a stroke, and are too slow to draw with. We present ProgressNet, a training-free framework that lets a frozen text-to-image model follow a drawing session as it unfolds: strokes are added and erased, the prompt is revised, and the image keeps up at about a second per turn. It needs no new parameters because the frozen model already has what a progressive generator needs, a pathway through which the previous turn can be remembered, layers that can carry appearance forward without freezing structure, and an internal signal of how far to trust an unfinished sketch; three inference-time mechanisms (Previous-Concept Memory, Layer-Selective K/V Injection and Banded Adaptive Control) use each in turn. As a sketch fills in, every existing method degrades, the FID of the FLUX+ControlNet baseline doubling between 10% and 100% completion on FS-COCO, while ProgressNet's barely moves; it maintains strong fidelity and progressive coherence across three sketch domains and is preferred by users over five competitors, most widely on erasure.