arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Gutenberg:利用近数据处理驯服延迟关键型云服务

Gutenberg: Taming Latency-Critical Cloud Services with Near-Data-Processing

Qiushi Lin, Phillip B. Gibbons, Jovan Stojkovic, Yiwei Zhao

arXiv 2609.06691首次发表:更新:

发表机构

University of Texas at Austin; Carnegie Mellon University(德克萨斯大学奥斯汀分校; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Gutenberg通过CPU+NDP架构、增量缓冲和在线控制器,降低延迟关键型云服务的尾延迟,提升隔离与公平性。

AI 中文摘要

延迟关键型云服务对内存施加了越来越大的压力,同时还需要隔离性、公平性和可预测的服务质量。近数据处理(NDP)通过在内存附近执行请求来减少数据移动,先前的系统通过缓存和复制进一步改善了局部性。然而,写入操作使得副本维护成本高昂,而不均衡的计算和内存流量可能会使少数NDP单元过载并增加尾延迟。现有的面向吞吐量的调度器无法完全解决这些共存云服务所面临的挑战。我们提出了Gutenberg,一种面向可变、延迟关键型云服务的CPU+NDP系统。Gutenberg在CPU驻留的增量缓冲区中暂存子页更新,使得热可写页面能够保持复制而无需急切的全页同步。它还采用CPU辅助核心,在NDP执行或副本维护成本高昂时协助请求执行。一个在线控制器利用访问模式、队列压力以及先前决策的反馈,联合决定页面放置、复制、CPU/NDP执行和路由。该系统还强制实施跨服务的隔离和公平资源分配。我们还对CPU-NDP协调协议进行了模型检查以确保正确性。我们使用ZSim(配备Ramulator校准的内存时序)在TailBench上进行了评估。在所评估的服务中,Gutenberg优于先前的系统,平均延迟和p99延迟分别降低了高达80.4%和85.8%。它还改善了隔离性和公平性,同时适应不断变化的工作负载行为。

英文摘要

Latency-critical cloud services place growing pressure on memory while requiring isolation, fairness, and predictable QoS. Near-data processing (NDP) reduces data movement by executing requests close to memory, and prior systems further improve locality through caching and replication. However, writes make replica maintenance expensive, while uneven compute and memory traffic can overload a few NDP units and increase tail latency. Existing throughput-oriented schedulers do not fully address these challenges for co-located cloud services. We present Gutenberg, a CPU+NDP for mutable, latency-critical cloud services. Gutenberg stages subpage updates in a CPU-resident delta buffer, allowing hot writable pages to remain replicated without eager full-page synchronization. It also adopts CPU helper cores to assist request execution when NDP execution or replica maintenance becomes costly. An online controller jointly decides page placement, replication, CPU/NDP execution, and routing using access patterns, queue pressure, and feedback from prior decisions. The system further enforces isolation and fair resource allocation across services. We also model-check CPU--NDP coordination protocol for correctness. We evaluate on TailBench using ZSim with Ramulator-calibrated memory timing. Across evaluated services, Gutenberg outperforms prior systems, reducing average and p99 latency by up to 80.4% and 85.8%. It also improves isolation and fairness while adapting to changing workload behaviors.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑