arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解释文本到图像扩散模型中对象依赖的概念脆弱性

Interpreting Object-Dependent Concept Brittleness in Text-to-Image Diffusion Models

Yifan Yuan, Xiangyu Liu, Hongming Shan, Yu Han, Yu Jiang, Hao Tan, Junping Zhang, Linlin Shen

arXiv 2609.09909首次发表:更新:

发表机构

Shenzhen University; Fudan University; National University of Singapore(深圳大学; 复旦大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文本到图像扩散模型中对象依赖的概念脆弱性,提出基于稀疏自编码器空间去噪轨迹分析的审计与推理时修正框架,通过原型插值显著提升概念一致性、文本保真度和修复成功率。

AI 中文摘要

尽管文本到图像扩散模型通常表现出强大的提示跟随能力,我们识别出一种持续存在且此前未被充分探索的失败模式,即在相同的生成设置下,仅对象不同的一小部分提示始终无法实现相同的目标概念。我们将此现象称为对象依赖的概念脆弱性。此类情况表明系统存在内部盲点,而非随机采样噪声。在本文中,我们提出了一个面向可解释性的框架来审计并最小化修正这些失败。我们的关键思想是在逐步稀疏自编码器(SAE)空间中分析去噪轨迹,在该空间中,抽象风格和属性概念比原始去噪表示更易分离。这种稀疏空间使我们能够比较成功和失败的生成,识别证据缺失、减弱或时间延迟的概念维度,并从可靠的类一致样本中构建类级概念原型。基于此审计过程,我们引入了一种轻量级的推理时修正策略,将去噪特征向SAE空间中对应的原型插值。这种干预并非作为特定任务的重训练方法,而是对所诊断概念缺陷的验证。我们在多个扩散骨干网络上评估了所提出的框架,针对风格和属性失败案例,在概念一致性、文本保真度和修复成功率方面均取得了显著改进。进一步的分析表明,更深层的去噪表示提供了更清晰的概念结构,而早期干预提供了最强的修正杠杆。代码可在以下网址获取:此 https URL。

英文摘要

Although text-to-image diffusion models generally exhibit strong prompt-following ability, we identify a persistent and previously underexplored failure pattern in which a small subset of prompts differing only in the object consistently fails to realize the same target concept under identical generation settings. We term this phenomenon object-dependent concept brittleness. Such cases suggest systematic internal blind spots rather than random sampling noise. In this paper, we present an interpretability-oriented framework to audit and minimally correct these failures. Our key idea is to analyze denoising trajectories in a step-wise sparse autoencoder (SAE) space, where abstract style and attribute concepts become more separable than in the raw denoising representation. This sparse space enables us to compare successful and failed generations, identify concept dimensions whose evidence is missing, weakened, or temporally delayed, and construct class-level concept prototypes from reliable class-consistent samples. Based on this audit process, we introduce a lightweight inference-time correction strategy that interpolates denoising features toward the corresponding prototype in SAE space. Rather than serving as a task-specific retraining method, this intervention acts as a validation of the diagnosed concept deficiency. We evaluate the proposed framework on style and attribute failure cases across multiple diffusion backbones, with significant improvements in concept consistency, text fidelity, and repair success. Further analyses show that deeper denoising representations provide clearer concept structure, while early-stage intervention offers the strongest correction leverage. Code is available at https://github.com/Metecade/Object-Dependent-Concept-Brittleness.

CommentsAccepted at ACM MM 2026. 27 pages, 17 figures, including appendices

DOI:10.1145/3767308.3835084

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑