发表机构
Not Community Labs Inc.; Intel Corporation(Not Community Labs公司; 英特尔公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Cascadia提出一种无控制平面的超融合AI基础设施,利用Intel AIPC的CPU、GPU和NPU资源,通过libp2p网格和CA证书实现去中心化调度,支持多种服务模式,显著提升响应吞吐量。
AI 中文摘要
我们提出了Cascadia,一个用于在商品化Intel AIPC机群上利用其CPU、集成GPU和NPU资源来服务大型语言模型的系统。每个节点都嵌入了入口、调度和执行功能;推理请求不需要专用的路由控制平面。节点使用CA签发的ed25519准入证书加入libp2p QUIC网格,通过直接对等流交换实时负载并传播签名能力,并将兼容OpenAI的请求路由到符合条件的对等节点。由运营商运行的证书颁发机构在推理路径之外处理准入和机群管理。三种服务模式共享一个接口:单节点上的整模型执行、负载均衡副本以及使用我们配套论文中的编译和推测解码机制的流水线分片链。可选的KV缓存移动性在路由移动后重用兼容的对话前缀,未命中时则进行冷重计算。签名响应收据和哈希链日志支持来源和审计。一个三节点Phi-3.5-mini NPU测试平台在十个并发请求下提供了其单节点配置响应吞吐量的3.10倍;另一个四节点部署记录了直接单节点服务吞吐量的4.06倍。成对的延迟观测、运行时测量和内部功能检查描述了所测试的配置。我们使用供应商文档,将Cascadia与IBM、Nutanix、VMware和HPE平台在部署占用空间、硬件要求、调度、扩展、许可和信任方面进行了比较。论文仓库提供了基准测试脚本、精选测量结果和声明到证据的映射。
英文摘要
We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh using CA-issued ed25519 admission certificates, gossip signed capabilities, exchange live load over direct peer streams, and route OpenAI-compatible requests to eligible peers. An operator-run certificate authority handles admission and fleet management outside the inference path. Three serving modes share one interface: whole-model execution on one node, load-balanced replicas, and pipeline-sharded chains using the compilation and speculative decoding mechanism of our companion paper. Optional KV-cache mobility reuses compatible conversation prefixes after a routing move, with cold recomputation on a miss. Signed response receipts and hash-chained logs support provenance and audit. A three-node Phi-3.5-mini NPU testbed delivered 3.10x the response throughput of its one-node configuration under ten concurrent requests; a separate four-node deployment recorded 4.06x the throughput of direct single-node serving. Paired latency observations, runtime measurements, and internal functional checks characterize the tested configurations. We compare Cascadia with IBM, Nutanix, VMware, and HPE platforms on deployment footprint, hardware requirements, scheduling, scaling, licensing, and trust, using vendor documentation. The paper repository provides benchmark scripts, curated measurements, and a claim-to-evidence map.
Comments26 pages