FFN压缩会改变下游什么?扩散语言模型中的同状态因果恢复
What does FFN compression change downstream? Same-state causal restoration in diffusion language models
浏览论文内容
中文总结 AI 辅助
针对扩散语言模型压缩中局部误差无法反映下游影响的问题,提出同状态因果恢复(SSR)方法,通过恢复关键FFN转换,在激进压缩下恢复89.9%精度并保留36.8%计算节省。
中文摘要 AI 辅助
扩散语言模型(DLMs)支持灵活、并行的生成,但其迭代去噪过程在计算上仍然昂贵,这促使人们采用日益激进的压缩策略。现有的压缩目标主要衡量压缩后的计算在局部上对原始计算的近似程度,但局部误差并不能揭示哪些被移除的计算对下游去噪轨迹真正重要。我们提出了同状态因果恢复(SSR),该方法在压缩模型到达的精确当前输入上恢复原始前馈网络(FFN),并衡量由此产生的轨迹变化。据我们所知,这是首次直接测量DLM压缩中移除FFN计算对同当前输入闭环效应的方法。在LLaDA-8B-Instruct和Dream-v0-Instruct-7B上,压缩侧状态排名在下游效应方面显著优于固定去噪阶段下的局部NMSE,而受控干预表明,校正结构比幅度更重要。使用无任务标签的校准,SSR为留出推理冻结了一个单一的恢复窗口。在激进的LLaDA压缩下,仅恢复四个转换即可恢复89.9%的丢失精度,同时保留估计的36.8%全模型MAC节省,并优于同等预算的局部误差基线。Dream进一步表明,恢复稠密行为和修复最终任务是不同的结果。
英文摘要
Diffusion language models (DLMs) enable flexible, parallel generation, but their iterative denoising remains computationally expensive, motivating increasingly aggressive compression. Existing compression objectives largely measure how well compressed computation approximates the original locally, but local error does not reveal which removed computations actually matter to the downstream denoising trajectory. We introduce Same-State Causal Restoration (SSR), which restores the original FFN on the exact current input reached by the compressed model and measures how the resulting trajectory changes. To our knowledge, this is the first direct measurement of the same-current-input closed-loop effect of removed FFN computation in DLM compression. Across LLaDA-8B-Instruct and Dream-v0-Instruct-7B, compressed-side state ranks this downstream effect substantially better than local NMSE at fixed denoising phase, while controlled interventions show that correction structure matters beyond magnitude. Using task-label-free calibration, SSR freezes a single restoration window for held-out inference. Under aggressive LLaDA compression, restoring only four transitions recovers 89.9% of the lost accuracy while retaining an estimated 36.8% whole-model MAC saving and outperforming an equal-budget local-error baseline. Dream further shows that restoring dense behavior and repairing the final task are distinct outcomes.