arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05635cs.DCcs.OS

DejaVu:统一内存分配以消除统一内存SoC上的冗余拷贝

Echo: Merging Host Device Buffers to Avoid Redundant Data Movement on Unified Memory SoCs

发表机构北卡罗来纳州立大学
查看机构详情
  • North Carolina State University(北卡罗来纳州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuheng Zhu, Yanbo Zhao, Jiajia Li, Man-Ki Yoon

首次发表
浏览论文内容

中文总结 AI 辅助

DejaVu通过编译时转换和二进制优化两条路径,在统一内存SoC上安全消除冗余CPU-GPU拷贝,拷贝密集型负载加速最高6.9倍,闭源应用加速1.10-1.14倍。

中文摘要 AI 辅助

统一内存(UMA)边缘平台上的GPU应用程序通常继承了离散GPU的内存抽象,在这种抽象中,它们为CPU分配一个缓冲区,为GPU分配另一个缓冲区,并在GPU执行前后在两者之间拷贝数据。在UMA硬件上,这些缓冲区位于同一物理DRAM中,因此这些拷贝消耗带宽、时间和能量,而并未跨越物理边界移动数据。尽管UMA平台的应用日益广泛,但这种模式仍然普遍存在,因为生产软件栈、库和示例是为跨离散GPU的可移植性而编写的。然而,消除这些拷贝并非简单地将两个缓冲区合并,因为原始程序可能依赖于两个缓冲区是不同的,或者依赖于拷贝本身来排序CPU和GPU的访问。DejaVu仅在程序不依赖上述效应时消除这些拷贝。它根据源代码是否可用,沿着两条互补的路径实现这一点。DejaVu-SR是一种编译时的LLVM转换,它证明安全性并就地重写可接受的配对。DejaVu-DR是一种面向闭源部署的基于profile的二进制优化器,它对稳定的分配/拷贝模式进行profile分析和验证,并在运行时拦截匹配的调用以合并profile匹配的配对,同时保留被移除拷贝的排序效应。在三个NVIDIA Jetson平台上的七个基准测试中,DejaVu的收益随着基线时间中拷贝所占比例的增加而增长。拷贝密集型工作负载的加速比最高达6.9倍,闭源端到端应用程序的加速比为1.10-1.14倍。源代码路径和二进制路径均能达到手动优化可实现性能的≥99%。

英文摘要

GPU applications on unified-memory (UMA) edge platforms often inherit a discrete-GPU memory abstraction in which they allocate one buffer for the CPU, another for the GPU, and copy data between them before and after GPU execution. On UMA hardware these buffers reside in the same physical DRAM pool, so the copies consume bandwidth, time, and energy. Despite the growing adoption of UMA platforms, this pattern remains common because production software stacks, libraries, and samples were written for portability across discrete GPUs. However, removing these copies is not as simple as merging the two buffers, because the original program may rely on the two buffers being distinct, or on the copy itself ordering CPU and GPU accesses. Echo removes these copies only when it preserves the data values and access ordering on which the program depends. It does so along two complementary paths: source-level rewriting and binary deployment. Echo-SR is a compile-time LLVM transformation that proves safety and rewrites accepted pairs in place. Echo-DR is a profile-guided binary optimizer for existing applications that profiles and validates stable allocation/copy patterns and, at runtime, intercepts the matching calls to unify profile-matched pairs while preserving the ordering effects of removed copies. Across seven benchmarks on three NVIDIA Jetson platforms, Echo's benefit grows with the fraction of baseline time spent on copies. Copy-dominated workloads speed up by up to 7.05x, closed-source end-to-end applications speed up by up to 1.40x. On the five source-available Orin benchmarks, both automatic paths recover >=99% of the gain of a manually optimized reference.

补充信息

↑