PulseQuant:面向4比特视频扩散Transformer的传播引导子空间校正
PulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion Transformers
浏览论文内容
中文总结 AI 辅助
PulseQuant提出结合轨迹敏感性与激活几何的4比特训练后量化方法,通过传播风险引导校准和子空间校正,在视频扩散Transformer上提升关键一致性指标。
中文摘要 AI 辅助
视频扩散Transformer中的量化误差会被后续的去噪更新放大或衰减,这使得局部重建误差成为最终影响的非完整预测指标。我们提出PulseQuant,一种4比特训练后量化方法,将轨迹敏感性与激活几何相结合以指导离线校准。孤立的块-步干预用于估计传播风险,该风险在行半径选择过程中优先考虑敏感轨迹状态。在固定这些半径后,响应子空间校正利用相邻码编辑来减少沿主导激活方向的残差分量。两个阶段均保留原始的4比特权重表示。受控干预表明,短时程传播误差比即时块输出误差更可靠地预测最终潜在误差,支持超越局部重建目标的校准。在Wan模型、Self Forcing和MiniMax-H3上的评估表明,关键一致性和密集参考指标得到改善,同时在不同模型规模和生成范式下保持其他属性的竞争力。
英文摘要
Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method that combines trajectory sensitivity with activation geometry to guide offline calibration. Isolated block--step interventions estimate propagation risk, which prioritizes sensitive trajectory states during row-radius selection. With these radii fixed, response-subspace correction uses neighboring-code edits to reduce residual components along dominant activation directions. Both stages preserve the original 4-bit weight representation. Controlled interventions show that short-horizon propagated error predicts final latent error more reliably than immediate block-output error, supporting calibration beyond local reconstruction objectives. Evaluations on Wan models, Self Forcing, and MiniMax-H3 demonstrate improvements in key consistency and dense-reference metrics while remaining competitive on other attributes across model scales and generation paradigms.
发表机构
- The University of Sydney(悉尼大学)
- The Hong Kong University of Science and Technology(香港科技大学)
- Shanghai AI Laboratory(上海人工智能实验室)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。