在掩码外锚定指令:用于高效上下文扩散Transformer的精确引用缓存
Beyond Attention Masks: Instruction Anchoring for Efficient In-Context Diffusion Generation
浏览论文内容
中文总结 AI 辅助
该研究针对上下文扩散Transformer中引用增多导致计算量过大的问题,提出掩码外锚定指令的精确引用缓存方法,通过静态文本锚点结合速度蒸馏实现高效图像编辑,在保持生成质量的同时大幅提升了去噪速度。
中文摘要 AI 辅助
多模态生成是各类内容创建与编辑应用的核心,上下文条件化是该范式的关键,它使扩散Transformer能在共享注意力序列中处理文本指令与视觉引用。然而,每张引用图像会引入数千个token,计算量随引用数量快速增长。现有方法通过结构化稀疏注意力减少计算,但该方法限制了引用与目标token间的交互,且使引用的键(K)和值(V)与去噪目标无关,可一次性计算并跨步骤复用,却会阻止视觉引用关注文本指令,大幅降低多引用编辑中的指令遵循度与引用保真度。为解决这一矛盾,我们联合重新设计token序列与注意力掩码,提出的掩码外设计使用静态文本锚点连接指令与引用分支,在不增加参数的情况下保留了精确的K和V复用。但这种直接架构转换会降低生成质量,我们通过教师强制速度蒸馏恢复丢失的性能,随后进入短的在线策略阶段,由教师监督学生访问的状态。据我们所知,这是首次将在线策略蒸馏用于扩散模型的架构恢复。在三个图像编辑基准上,我们的方法达到了全注意力的生成质量;使用5张引用图像时,它将完整的40步去噪过程加速了3.92倍,而静态文本锚点引入的运行时开销可忽略;在缩放研究中,10张引用时的加速比达到5.47倍。
英文摘要
In-context diffusion transformers concatenate instruction, target, and reference tokens into a single sequence for joint attention. Reference-side computation must therefore be repeated at every denoising step, with the cost growing rapidly as more references are added. Decoupling reference tokens from the target enables exact key-value reuse across denoising steps, but prevents the references from attending to the instruction, degrading instruction following and reference fidelity. This trade-off cannot be resolved through attention-mask design alone. We introduce AnchorCache, a parameter-free token-layout and attention-mask co-design that inserts static text anchors. These anchors condition the reference representations on the instruction during cache construction, after which the resulting reference keys and values can be reused exactly across denoising steps. To recover the quality initially lost through this structural conversion, we apply teacher-forced velocity distillation followed by a short on-policy stage that queries the teacher at student-visited states. To our knowledge, this is the first use of on-policy distillation for architectural recovery in diffusion models. Across benchmarks spanning image, speech, and video generation, AnchorCache matches full-attention quality. Its efficiency gains increase with the reference-context size, reaching a 6.40x speedup in diffusion transformer inference.
发表机构
- Harbin Institute of Technology(哈尔滨工业大学)
- KlingAI Research(KlingAI研究院)
机构由 AI 辅助整理,请以论文原文为准。