arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15569cs.OS

使用基于弹性CXL的分布式共享内存扩展未修改的多线程应用程序

Scaling Unmodified Multithreaded Applications with Elastic CXL-based Distributed Shared Memory

Guowei Liu, Kang Chen, Laiping Zhao, Yiming Li, Hanwen Liu, Chen Peng, Yichi Chen, Sheng Chen, Zhiyuan Su, Wenyu Qu

首次发表
浏览论文内容

中文总结 AI 辅助

研究如何用基于弹性CXL的分布式共享内存扩展未修改的多线程应用程序。核心方法是提出xDSM系统,采用操作系统-运行时协同设计、动态延迟驱动策略和空间局部性感知弹性。主要贡献是性能比基线和现有混合DSM有显著提升,且实现近线性可扩展性。

中文摘要 AI 辅助

虽然CXL为分布式共享内存(DSM)提供了一个有前景的硬件基础,但在多个节点间无缝扩展多线程应用程序仍然是一个巨大挑战。现有的基于CXL的DSM存在不足,如需要手动代码修改来共享非堆数据、采用的刚性数据放置策略在多样动态工作负载下失效、在亚微秒($\mu\mathrm{s}$)环境中存在严重的页面错误处理开销。我们提出了xDSM,这是一个基于CXL构建的全空间、弹性DSM系统,能透明地扩展未修改的多线程应用程序。它采用操作系统-运行时协同设计来消除手动代码重写负担,建立全局协调地址空间以无缝共享所有内存段;放弃静态放置规则,采用动态、延迟驱动策略在本地DRAM和CXL内存间主动平衡数据;引入空间局部性感知弹性,动态合并和拆分页面以分摊处理成本。在15种系统配置下对多种工作负载进行评估,xDSM比仅使用CXL的基线性能高出1.5倍至2.2倍,比最先进的混合DSM高出1.1倍至2.2倍,同时实现了近线性可扩展性。

英文摘要

While CXL presents a promising hardware substrate for Distributed Shared Memory (DSM), seamlessly scaling multithreaded applications across multiple nodes remains a formidable challenge. Existing CXL-based DSMs fall short: they require manual code modifications to share non-heap data, employ rigid data placement policies that fail under diverse and dynamic workloads, and suffer from severe page-fault processing overheads in sub-microsecond ($μ\mathrm{s}$) environments. We present xDSM, a full-space, elastic DSM system built over CXL that transparently scales unmodified multithreaded applications. To eliminate the burden of manual code rewrites, xDSM employs an OS-runtime co-design that establishes a globally coordinated address space, seamlessly sharing all memory segments. To mask CXL access penalties, xDSM abandons static placement rules in favor of a dynamic, latency-driven policy that actively balances data between local DRAM and CXL memory. Finally, to resolve the fundamental tension between high base-page fault overheads and severe huge-page false sharing, xDSM introduces spatial locality-aware elasticity, dynamically coalescing and splitting pages on the fly to amortize processing costs. Evaluated across diverse workloads using 15 system configurations, xDSM outperforms CXL-only baselines by 1.5$\times$ to 2.2$\times$ and state-of-the-art hybrid DSMs by 1.1$\times$ to 2.2$\times$, while achieving near-linear scalability.

补充信息

↑