arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14789cs.AR

瓦利诺:对快速、节能且可编程的物理内存分配的架构支持

Valinor: Architectural Support for Fast, Energy-Efficient and Programmable Physical Memory Allocation

Konstantinos Kanellopoulos, Spiros Galanopoulos, Konstantinos Sgouras, Vlad-Petru Nitu, Ilias Papalamprou, Andreas Kosmas Kakolyris, Rahul Bera, Dimosthenis Mas… 展开作者

Konstantinos Kanellopoulos, Spiros Galanopoulos, Konstantinos Sgouras, Vlad-Petru Nitu, Ilias Papalamprou, Andreas Kosmas Kakolyris, Rahul Bera, Dimosthenis Masouros, Dimitrios Soudris, Onur Mutlu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对当前系统物理内存分配开销大的问题,提出瓦利诺这一硬件 - 操作系统协作的内存分配基板,通过可编程硬件分配引擎结合多种策略,在实际硬件和模拟中实现加速分配、提升性能及降低能耗等效果。

中文摘要 AI 辅助

物理内存分配按需建立虚拟到物理的映射。在当前系统中,每个小页面错误都会陷入内核并触发流水线刷新、停顿以及一系列可能花费数万个周期的分配步骤。对于无服务器函数和微服务等短期工作负载,这些开销愈发显著。先前的硬件分配提议存在不足。我们提出了瓦利诺,一种硬件 - 操作系统协作的内存分配基板,它结合了软件灵活性与硬件级性能。瓦利诺引入了可编程硬件分配引擎,支持多种策略。我们在运行Linux的BOOM RISC - V软核和全系统模拟器上实现了瓦利诺。在实际硬件上,它加速分配17倍,提高端到端性能16%,降低能耗达8%。全系统模拟进一步评估了可编程分配引擎和六个分配库,表明瓦利诺在不牺牲可编程性的情况下提供硬件级性能。

英文摘要

Physical memory allocation establishes virtual-to-physical mappings on demand. In current systems, each minor page fault traps into the kernel and triggers pipeline flushes, stalls, and a long sequence of allocation steps that can cost tens of thousands of cycles. These overheads are increasingly significant for short-lived workloads such as serverless functions and microservices, where minor faults can account for up to 54% of runtime and up to 40% of system energy. Prior hardware allocation proposals avoid traps and context switches, but either sacrifice useful placement optimizations or rely on fixed-function logic that cannot adapt to new policies or changing hardware conditions. We present Valinor, a hardware-OS cooperative memory allocation substrate that combines software flexibility with hardware-class performance. Valinor introduces a programmable hardware allocation engine that executes compact OS-supplied allocation libraries at close to fixed-hardware speed. It supports diverse policies, including short-lived object allocators, integrity mechanisms, and hardware-telemetry-guided placement. We implement Valinor on a BOOM RISC-V soft core running Linux and in a full-system simulator. On real hardware, Valinor accelerates allocation by 17x, improves end-to-end performance by 16%, and reduces energy consumption by up to 8%. Full-system simulation further evaluates the programmable allocation engine and six allocation libraries, showing that Valinor provides hardware-class performance without sacrificing programmability.

↑