arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Cascadia:在十一台 AI PC 上常驻运行 975B 参数的 MoE 推理

Cascadia: Resident 975B MoE Inference on Eleven AI PCs

Tate Berenbaum, Matias Parij, Muthaiah Venkatachalam

arXiv 2610.07219首次发表:更新:

发表机构

Not Community Labs Inc.; Intel Corporation(Not Community Labs 公司; 英特尔公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 Cascadia 系统,在十一台 AI PC 上常驻运行 975B 参数 MoE 模型,通过定制引擎优化推理,实现 60.29 tokens/s 解码速度,并验证了长上下文下的代码恢复能力。

AI 中文摘要

混合专家模型通过稀疏的每词元计算,使得近万亿参数的容量变得可访问,前提是服务系统能够分配权重并协调其执行。我们展示了 Cascadia 对 Inkling 的常驻执行,这是一个总参数 975B、激活参数 41B 的模型,运行在十一台 Intel Core Ultra X7 358H AI PC 上,每台配备 64 GB 内存、Arc B390 集成显卡和千兆以太网。我们贡献了一个定制的常驻 MoE 引擎,该引擎保留了 Inkling 的路由规则,为 OpenVINO 的融合 iGPU 原语构建压缩图,并协调 FP16 专家计算与 FP32 输出恢复。该引擎每台机器容纳六个连续的解码器层,并将密集前馈块表示为全激活专家切片,将实测的密集层调用时间从约 8.1 毫秒减少到 4.5 毫秒。一个流式流水线协调并发生成,而捕获状态的草稿评估衡量与部署数值路径的一致性。在从 1 到 176 个流的十五个并发级别上的成对测量,在 88 个流时达到 60.29 个聚合解码词元/秒,在整个服务阶段为 46.87 个词元/秒。在十五个流时,中位首词元延迟为 6.05 秒。将上下文预算从默认的 1,024 位置提高后,1k 到 64k 词元的真实提示在所有 19 个测量答案中恢复了嵌入代码,首词元时间随 $aN+bN^2$ 增长,解码延迟近似线性增长,两者均受单线程 CPU 注意力循环限制,而非内存限制,内存每流可容纳 512k 位置。对捕获的舰队状态的评估分离了词汇选择和权重量化对草稿一致性的影响。这些贡献共同为分布式客户端系统上具有共享 CPU-GPU 内存的大型稀疏模型建立了一种执行和评估方法。

英文摘要

Mixture-of-experts models make nearly trillion-parameter capacity accessible with sparse per-token computation, provided that the serving system can distribute the weights and coordinate their execution. We present Cascadia's resident execution of Inkling, a 975B-total/41B-active-parameter model, on eleven Intel Core Ultra X7 358H AI PCs, each with 64 GB of memory, Arc B390 integrated graphics and gigabit Ethernet. We contribute a custom resident MoE engine that preserves Inkling's routing rules, constructs compressed graphs for OpenVINO's fused iGPU primitives, and coordinates FP16 expert computation with FP32 output restoration. The engine fits six consecutive decoder layers per machine and represents dense feed-forward blocks as all-active expert slices, reducing measured dense-layer call time from approximately 8.1 to 4.5 ms. A streaming pipeline coordinates concurrent generation, while captured-state draft evaluation measures agreement with the deployed numerical path. Paired measurements at fifteen concurrency levels from 1 to 176 streams reach 60.29 aggregate decode tokens/s at 88 streams, with 46.87 tokens/s over the complete serving phases. At fifteen streams, median first-token latency is 6.05 s. Raising the context budget from the 1,024-position default, real prompts of 1k to 64k tokens recover the embedded code in all 19 measured answers, with first-token time growing as $aN+bN^2$ and decode latency growing approximately linearly, both bounded by a single-threaded CPU attention loop rather than by memory, which holds 512k positions per stream. Evaluation on captured fleet states separates the effects of vocabulary selection and weight quantization on draft agreement. Together, these contributions establish an execution and evaluation approach for large sparse models on distributed client systems with shared CPU-GPU memory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑