时间分辨的 Token 归因揭示扩散语言模型的生成动态
Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
本文提出扩散层积分梯度(DLIG),一种用于扩散语言模型的 token 归因方法,通过扩展积分梯度至任意层和去噪步骤,揭示模型生成动态,并在词义消歧、多跳推理和句子填充任务中验证其有效性。
中文摘要 AI 辅助
本文提出了扩散层积分梯度(DLIG),一种针对扩散语言模型(DLM)的 token 归因方法,它将积分梯度(IG)扩展到任意层和去噪步骤。DLIG 将 DLM 对自生成或固定补全的渐进承诺归因于输入提示。我们建立了 DLIG 与 IG 公理(完整性、实现不变性、线性性和对称性保持)之间的直接对应关系。作为干预分析的轻量级补充,DLIG 提供了一种廉价的初步检查机制,用于跨去噪轨迹的机制假设。我们在词义消歧、多跳图推理和句子填充任务上进行了演示,揭示了 DLM 如何跨位置、层和去噪步骤利用输入信息。
英文摘要
This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.
发表机构
- Université Grenoble Alpes(格勒诺布尔阿尔卑斯大学)
- CentraleSupélec(巴黎中央理工-高等电力学院)
机构由 AI 辅助整理,请以论文原文为准。