AI 中文总结
研究在边缘设备部署视频扩散模型时速度慢的问题,提出CODA架构,通过计算-缓存算子分解,分离计算与缓存路径、重组缓存活动并利用分支独立性,实现端到端加速和能源效率提升,且保持生成质量。
AI 中文摘要
在边缘设备上部署视频扩散模型(VDM)对局部和隐私保护生成很有吸引力,但基于Transformer的迭代去噪对于实际的局部推理来说仍然太慢。跨时间步长缓存(CTC)已成为减少冗余计算的一个有前途的方向,它在相邻去噪步骤之间重用激活而不是修改模型权重,同时很大程度上保留生成保真度。然而,在内存受限的边缘GPU上,CTC需要大量的缓存占用,这很快就会超过设备上的VRAM并迫使缓存进入主机内存。更根本的是,缓存算子仍然与原生计算算子紧密交织且具有链依赖性,所以简单的近内存卸载仍然会因残差和融合计算而产生重复的PCIe交换,使缓存重用变成一个受通信和序列化限制的执行流程。因此,我们提出了CODA,一种以计算-缓存算子分解为中心的算法-硬件协同设计架构。CODA在xPU和轻量级DIMM侧近内存引擎之间分离密集计算路径和内存受限的缓存路径,将碎片化的缓存活动重新组织成硬件友好的合并段,并利用无分类器指导(CFG)分支独立性使xPU计算与缓存侧执行重叠。实验表明,CODA实现了高达1.80倍的端到端加速和1.74倍的更高能源效率,同时与最先进的缓存算法相比保持了有竞争力的生成质量。
英文摘要
Deploying Video Diffusion Models (VDMs) on edge devices is appealing for localized and privacy-preserving generation, but their iterative Transformer-based denoising remains too slow for practical local inference. Cross-Timestep Caching (CTC) has emerged as a promising direction for reducing redundant computation, reusing activations across adjacent denoising steps rather than modifying model weights, while largely preserving generation fidelity. However, on memory-constrained edge GPUs, CTC requires a massive cache footprint that quickly exceeds on-device VRAM and forces the cache into host memory. More fundamentally, cache operators remain tightly interleaved and chain-dependent with native compute operators, so naive near-memory offloading still incurs repeated PCIe exchanges for residual and fusion computations, turning cache reuse into a communication- and serialization-bound execution flow. We therefore propose CODA, an algorithm-hardware co-designed architecture centered on Compute-Cache Operator Disaggregation. CODA separates dense compute paths and memory-bound cache paths across the xPU and a lightweight DIMM-side near-memory engine, reorganizes fragmented cache activity into hardware-friendly coalesced segments, and exploits Classifier-Free Guidance (CFG) branch independence to overlap xPU compute with cache-side execution. Experiments show that CODA achieves up to 1.80x end-to-end speedup and 1.74x higher energy efficiency, while preserving competitive generation quality compared with a state-of-the-art caching algorithm.
Comments15 pages, 14 figures, accepted to the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)