AI 中文总结
该研究提出 fabric_ext 中间件编译器与运行时,通过语义移动图实现跨 GPU-CXL Fabric 多层的 eBPF 策略,以处理 LLM 预填充等场景下的跨岛数据流需求。
AI 中文摘要
我们提出 fabric_ext,这是一款用于 GPU-CXL Fabric 上可扩展操作系统策略的 eBPF 中间件编译器与运行时。fabric_ext 支持单个策略程序在 GPU 钩子、驱动/运行时钩子、DPU/NIC 钩子以及 CXL 交换机或近内存钩子间执行。其核心抽象为语义移动图:边描述字节、步长、复用距离、读写比例、源与目的地、排序要求、别名集、所有权,以及移动、量化、压缩、校验和、过滤、归约、散集/聚散、复制、持久化等变换。编译器将该图转换为各设备的 eBPF 程序、验证器义务、一致性分类的 BPF 映射,以及 bpftime 和 dputime 的构件。在 Fabric 边缘,fabric_ext 将近 Type-2 小型核心视为硬件 JIT 与状态管理器:它将已验证的移动描述符特化为本地复制、放置、排序和变换命令,而周围的冯·诺依曼内存、DMA 与计算引擎岛执行数据流。由于该数据流是数据驱动的,fabric_ext 还在岛旁放置观测点,使队列、DMA 完成、内存放置和所有权转换在发生时即可被观测。典型压力场景为 LLM 预填充:注意力流处理 KV 块与归约,而 FFN 流处理激活值、权重和可压缩中间件,迫使单个请求跨 GPU 张量执行、DPU/NIC 事件执行以及 CXL 或交换机本地数据流岛。
英文摘要
We present fabric_ext, an eBPF middleware compiler and runtime for extensible OS policies over GPU--CXL fabrics. fabric_ext lets one policy program execute across GPU hooks, driver/runtime hooks, DPU/NIC hooks, and CXL switch or near-memory hooks. The key abstraction is a semantic movement graph: edges describe bytes, stride, reuse distance, read/write ratio, source and destination, ordering requirement, alias set, ownership, and transformations such as Move, Quantize, Compress, Checksum, Filter, Reduce, Scatter/Gather, Replicate, and Persist. The compiler lowers this graph into per-device eBPF programs, verifier obligations, consistency-classed BPF maps, and artifacts for bpftime and dputime. At the fabric edge, fabric_ext treats a near-Type-2 small core as a hardware-JIT and state manager: it specializes verified movement descriptors into local copy, placement, ordering, and transformation commands, while the surrounding Von Neumann island of memory, DMA, and compute engines performs the dataflow. Because this dataflow is data-driven, fabric_ext also places observation beside the island, where queues, DMA completions, memory placement, and ownership transitions are visible as they happen. The canonical stress case is LLM prefill: attention streams KV blocks and reductions while FFN streams activations, weights, and compressible intermediates, forcing one request to cross GPU tensor execution, DPU/NIC event execution, and CXL or switch-local dataflow islands.