锚定与引导扩散:在推理时提升文生图生成的忠实度
Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
浏览论文内容
中文总结 AI 辅助
提出无训练框架AnchorSteer,通过语义锚定与反射引导组件改进文生图扩散模型的推理,在GenEval等基准上提升文本-图像对齐性能且保持视觉质量。
中文摘要 AI 辅助
尽管文本到图像的扩散模型能实现令人印象深刻的视觉质量,但它们常难以与复杂的组合提示保持精确对齐。一种有效策略是改进扩散模型的推理过程,从而更好地利用其预训练先验来解决对齐问题。现有的无训练方法可分为两类:第一类聚焦于改进随机采样的初始噪声,要么对噪声池进行代价高昂的搜索,要么在不确保可靠语义注入的情况下操纵采样噪声;第二类聚焦于改进去噪轨迹,缺乏明确的机制来及时诊断和纠正语义错误。我们提出AnchorSteer,这是一个无训练框架,可对初始化和去噪轨迹两者施加细粒度控制。AnchorSteer由两个协同组件构成:语义锚定(Semantic Anchoring)通过基于CLIP的先验提取和一种新颖的潜在先验分数蒸馏采样(LP-SDS)目标,将无信息的高斯噪声替换为与文本对齐的初始化;具体而言,LP-SDS将CLIP视觉先验蒸馏到扩散模型的知识分布中,以减轻基于CLIP的先验与基于扩散的先验之间的领域差距。反射引导(Reflective Steering)通过主动的思考-擦除-重绘(Think--Erase--Retouch)循环转变被动去噪,该循环支持生成中的自校正;它利用基于视觉语言模型(VLM)的诊断来检测语义偏差,并执行针对性的潜在细化以抑制错误内容并恢复缺失属性。在GenEval和T2I-CompBench++上的大量实验表明,AnchorSteer在保持高视觉质量的同时,在文本-图像对齐方面始终优于现有基线。
英文摘要
While text-to-image diffusion models achieve impressive visual quality, they frequently struggle to maintain precise alignment with complex compositional prompts. An effective strategy is to improve the inference process of diffusion models, thereby better leveraging their pretrained priors to address misalignment. Existing training-free methods can be divided into two categories. The first category focuses on improving the randomly sampled initial noise, either performing costly search over noise pools or manipulating sampled noise without ensuring reliable semantic injection. The second category focuses on improving the denoising trajectory, lacking explicit mechanisms to timely diagnose and correct semantic errors. we propose \textbf{AnchorSteer}, a training-free framework that exerts fine-grained control over \textbf{both initialization} and \textbf{the denoising trajectory}. AnchorSteer consists of two synergistic components: \textbf{Semantic Anchoring} replaces uninformative Gaussian noise with text-aligned initializations via CLIP-based prior extraction and a novel Latent-Prior Score Distillation Sampling (LP-SDS) objective. Specifically, LP-SDS distills CLIP visual priors into the knowledge distribution of diffusion models, mitigating the domain gap between CLIP-based priors and diffusion-based priors. \textbf{Reflective Steering} transforms passive denoising with an active Think--Erase--Retouch loop that enables mid-generation self-correction. It leverages VLM-based diagnosis to detect semantic deviations and performs targeted latent refinement to suppress erroneous content and recover missing attributes. Extensive experiments on GenEval and T2I-CompBench++ demonstrate that AnchorSteer consistently outperforms existing baselines in text--image alignment while preserving high visual quality.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。