发表机构
Max Planck Institute for Informatics; Sapienza University of Rome(马克斯·普朗克信息学研究所; 罗马大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本到图像流模型输出控制难题,提出Steering Fields自适应向量场,在生成轨迹上动态调整引导方向,实现安全引导与结构保持的图像编辑,并在基准上达到最先进水平。
AI 中文摘要
随着最先进的文本到图像流模型达到接近照片级真实感的质量,控制其输出(例如,抑制有害内容同时促进良性替代)已成为一个核心挑战。当前的引导范式包括向选定的激活添加一个全局引导向量。虽然这种方法有效,但一个固定且与示例无关的向量被均匀地应用于整个轨迹,无法适应生成过程中的变化状态,并且常常导致意外的全局变化。我们引入了Steering Fields,这是引导向量的一种推广,它在生成过程的每一步自适应地重新估计引导方向。Steering Fields作用于流模型的噪声状态,展现了引导强度与内容保留之间的连续权衡,并且是可组合的,能够同时诱导和抑制概念,在安全引导基准上达到了新的最先进水平。尽管没有使用显式的空间掩码或对象先验,轨迹自适应估计自然地保留了局部结构,这种方式类似于图像编辑。实际上,Steering Fields可以作为一种保持结构的图像编辑技术,在语义保真度(CLIP、VQAScore)上达到最先进水平,同时保持模型无关且无需反演。
英文摘要
As state-of-the-art text-to-image flow models achieve near-photorealistic quality, controlling their outputs, e.g., suppressing harmful content while promoting benign alternatives, has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected activations. While functional, a fixed and example-agnostic vector applied uniformly along the entire trajectory cannot adapt to the changing state of the generation and often causes unintended global changes. We introduce Steering Fields, a generalization of steering vectors that adaptively re-estimates the steering direction at each step of the generative process. Steering Fields operate on the noisy states of flow models, expose a continuous trade-off between steering strength and content preservation, and are compositional, enabling the simultaneous induction and inhibition of concepts, setting a new state of the art on safety steering benchmarks. Despite using no explicit spatial masks or object priors, the trajectory-adaptive estimation naturally preserves local structure, in a manner reminiscent of image editing. In fact, Steering Fields can serve as a structure-preserving image-editing technique that achieves state-of-the-art semantic fidelity (CLIP, VQAScore), while remaining model-agnostic and inversion-free.