发表机构
Key Lab of High Confidence Software Technologies (Peking University); State Key Laboratory of Networking and Switching Technology (BUPT)(高可信软件技术教育部重点实验室(北京大学); 网络与交换技术国家重点实验室(北京邮电大学))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OmniTide通过算法与系统协同设计,利用模态感知稀疏性,实现端侧全模态流式推理的高效与实时无限上下文处理。
AI 中文摘要
端侧流式全模态推理保障了用户隐私,并消除了高昂的按令牌API成本,但面临一个关键瓶颈:多模态数据的持续涌入,通过单调增长的KV缓存,迅速耗尽受限的内存和计算预算。现有的稀疏注意力方法存在不足,要么产生高昂的在线估计延迟,要么破坏交错的多模态上下文,同时未能解决物理内存碎片问题。我们提出OmniTide,这是首个专为高效端侧流式全模态推理设计的算法-系统协同方案。受模态感知的结构性稀疏性观察驱动,OmniTide采用基于单元(unit)的抽象,包含两个组件:(1)在算法层面,OmniPick基于单元边界和模态重要性,逻辑上保留关键的多模态上下文,以保持任务准确性;(2)在系统层面,OmniPage按保留可能性对缓存进行物理分区,并动态压缩存活的稀疏令牌,从而最小化内存碎片和数据移动开销。在三个流式基准和两种消费级设备架构上的广泛评估表明,OmniTide实现了高达12.72倍的核函数加速和2.40倍的流循环延迟降低。在StreamingBench上,与滑动窗口基线相比,在相当的会话成本下,其准确率提升了最多18.0个百分点。OmniPage相对于原生逻辑驱逐,进一步将物理KV跨度减少了最多26.7%,从而在边缘设备上实现了实时、无限上下文的流式处理。
英文摘要
On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention methods fall short, either incurring prohibitive online estimation latency or destroying interleaved cross-modal context, while failing to resolve physical memory fragmentation. We present OmniTide, the first algorithm-system co-design tailored for efficient on-device streaming omni-modal inference. Driven by the observation of modality-aware structural sparsity, OmniTide adopts a unit-based abstraction with two components: (1) At the algorithm level, OmniPick logically retains critical multimodal context based on unit boundaries and modality importance to preserve task accuracy; (2) At the system level, OmniPage physically partitions the cache by retention likelihood and dynamically compacts surviving sparse tokens, minimizing both memory fragmentation and data-movement overhead. Extensive evaluations across three streaming benchmarks and two consumer-device architectures show that OmniTide achieves up to $12.72\times$ kernel speedups and $2.40\times$ lower stream-loop latency. On StreamingBench, it improves accuracy by up to 18.0 percentage points over sliding-window baselines at comparable session cost. OmniPage further reduces the physical KV span by up to 26.7% relative to native logical eviction, unlocking real-time, infinite-context streaming on edge devices.