UniCache:面向统一多模态模型的任务与类型感知KV缓存压缩
UniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal Models
AI总结:
针对统一多模态模型中KV缓存压缩忽视任务与类型差异的问题,提出无需训练的UniCache框架,通过离线校准和注意力引导分配实现任务感知压缩,在理解、编辑和生成任务中分别达到5倍和2.5倍压缩,吞吐量提升1.78倍。
AI中文摘要:
统一多模态模型在单一网络内结合了理解、生成和编辑功能,为多功能多模态应用提供了有前景的基础。然而,不断增长的多模态上下文使得KV缓存的存储和访问成本日益高昂。现有的KV缓存压缩方法通常针对特定任务和单模态缓存进行定制,而忽视了跨任务和时间步的缓存重要性变化。然而,在统一多模态模型中,每个任务涉及多种KV缓存类型,且它们的组成和动态在不同任务间存在差异。因此,单一的压缩策略忽视了任务和类型特定的需求,导致关键信息丢失和跨任务质量下降。基于这些发现,我们提出了UniCache,一个无需训练的任务与类型感知KV缓存压缩框架。UniCache识别每个任务激活的缓存段,并通过离线校准分配合适的压缩策略。它通过注意力引导分配和任务感知时间调度,在共享存储预算下协调它们的并行执行。实验表明,UniCache在理解和编辑任务中实现了5倍的KV缓存压缩,在生成任务中实现了2.5倍的压缩,且质量损失可忽略不计,同时在长上下文设置中吞吐量提升高达1.78倍,显著提高了将统一多模态模型扩展到更长上下文的实用性。
英文摘要:
Unified multimodal models combine understanding, generation, and editing within a single network, offering a promising foundation for versatile multimodal applications. However, growing multimodal contexts make KV cache storage and access increasingly costly. Existing KV cache compression methods are typically tailored to specific tasks and single-modality caches, while overlooking changes in cache importance across tasks and timesteps. However, in unified multimodal models, each task involves multiple KV cache types, and both their composition and dynamics differ across tasks. As a result, a single compression policy overlooks task- and type-specific requirements, leading to the loss of critical information and degraded quality across tasks. Based on these findings, we propose UniCache, a training-free framework for task- and type-aware KV cache compression. UniCache identifies the cache segments activated by each task and assigns suitable compression policies through offline calibration. It coordinates their parallel execution under a shared storage budget through attention-guided allocation and task-aware temporal scheduling. Experiments show that UniCache achieves $5\times$ KV cache compression for understanding and editing and $2.5\times$ for generation with negligible quality loss, while increasing throughput by up to $1.78\times$ in long-context settings, significantly improving the practicality of scaling unified multimodal models to longer context.