D-Scope:利用稀疏自编码器分解和引导扩散Transformer
D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders
- Pivotal Research
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
- Schmidt Sciences(施密特科学)
- Stanford University(斯坦福大学)
- University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
D-Scope通过共享视觉证据将扩散Transformer的稀疏自编码器特征解释与生成控制结合,实现无需文本注释的特征检索和干预评估,并验证了对比检索在引导中的优势。
AI中文摘要:
稀疏自编码器(SAEs)揭示了扩散Transformer(DiTs)中的视觉结构,但解释一个特征并不能确定它是否可用于控制生成。我们引入了D-Scope(Diffusion Scope),这是一个通过共享视觉证据将特征解释与生成控制联系起来的框架。D-Scope将高激活图像块的SigLIP~2嵌入聚合成视觉质心。在共享的图像-文本嵌入空间中,将目标文本描述与这些视觉质心匹配,即可无需逐特征文本注释而检索单个特征。底层图像块为检查每个选择提供了证据,而空间掩蔽干预在固定生成条件下以不同强度测试相应的解码器方向。我们表征了两个模型家族和五个层中的150个SAEs,并引入了一个包含100个目标概念、每个概念十个上下文(涵盖欠指定和显式冲突条件)的基准。我们的实证结果表明,高重建保真度可以与低字典利用率和有限的视觉证据覆盖共存。在逐案例最佳扫描强度选择下,对比检索在测试的引导配置中比直接检索产生更大的平均区域SigLIP~2增益,但并未一致改善区域外保持。D-Scope提供了一个可检查的框架,通过视觉证据及其解码器方向对生成的影响来评估稀疏DiT特征。演示可在https://this https URL获取。
英文摘要:
Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shared visual evidence. D-Scope aggregates SigLIP~2 embeddings of highly activating image patches into visual centroids. Matching target text descriptions against these visual centroids in the shared image-text embedding space then enables retrieval of individual features without per-feature text annotations. The underlying patches provide evidence for inspecting each selection, while spatially masked interventions test the corresponding decoder direction at varying strengths under fixed generation conditions. We characterize 150 SAEs across two model families and five layers, and introduce a benchmark of 100 target concepts with ten contexts each spanning under-specified and explicit-conflict conditions. Our empirical results show that high reconstruction fidelity can coexist with low dictionary utilization and limited visual-evidence coverage. Under per-case best-of-sweep strength selection, contrastive retrieval yields larger mean regional SigLIP~2 gains than direct retrieval across the tested steering configurations, without consistently improving outside-region preservation. D-Scope provides an inspectable framework for evaluating sparse DiT features through their visual evidence and the effects of their decoder directions on generation. The demo is available at https://jiahaozhang-public.github.io/d-scope/.