arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19169cs.ARcs.DCcs.SYeess.SY

SiliconBench:统一内存桌面端LLM服务的速度、内存与保真度评估

SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops

发表机构宾夕法尼亚州立大学 · 伊利诺伊大学厄巴纳-香槟分校 · EXO实验室
查看机构详情
  • Penn State University(宾夕法尼亚州立大学)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • EXO Labs(EXO实验室)

机构由 AI 辅助整理,请以论文原文为准。

Ranran Haoran Zhang, Aysa Xuemo Fan, David Munhá Correia, Alex Cheema, Rui Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对统一内存桌面端并发LLM服务,提出SiliconBench框架,从速度、内存、保真度三维度评估九个引擎,发现vllm-metal吞吐量优势及内存预算局限,并发布基准代码与结果。

中文摘要 AI 辅助

在统一内存桌面端进行并发本地LLM服务时,必须保留内存余量并保证输出保真度,而仅关注速度的排名会忽略这些方面。我们提出了SiliconBench,它通过三个维度评估九个Apple Silicon服务引擎:速度、内存和保真度。我们在Qwen3、Qwen3.5和Gemma 4上评估了聊天和智能体服务。我们使用一个分类任务来检查相对于NVIDIA参考实现的质量回退。DGX Spark提供了补充的服务性能参考。三个期望指导了我们的解读:服务架构就绪度、内存纪律和多节点扩展。在Qwen3-0.6B上,仅vllm-metal在并发数从1到16时,两种工作负载的吞吐量都提高了一倍以上。CUDA vLLM和SGLang在相同提示上表现出更强的并发扩展性。显式内存预算并不能保证内存余量:两个栈完成了所有请求,但内存使用接近物理容量且吞吐量下降。较新的模型架构对引擎的支持范围更窄。它们评估的实现与保真度参考相匹配。只有三个栈满足完成度、保真度和模型覆盖的门槛。在更大密集型和MoE模型上的比较强化了在持续生成过程中调度提示处理的重要性:vllm-metal的打包预填充-解码路径在并发负载下保持了比omlx更低的首令牌延迟。在测试的双机配置中,基于Thunderbolt RDMA的张量并行扩展良好,而基于TCP的流水线并行则出现回退。我们发布了基准代码、每次运行结果和维护日志,并辅以结合受限智能体修复与人工审查的工作流程。

英文摘要

Concurrent local LLM serving on unified-memory desktops must preserve memory headroom and output fidelity, which speed-only rankings overlook. We introduce SiliconBench, which evaluates nine Apple Silicon serving engines through three lenses: speed, memory, and fidelity. We evaluate chat and agent serving on Qwen3, Qwen3.5, and Gemma 4. We use a classification task to check for quality regressions against an NVIDIA reference. DGX Spark provides a complementary serving-performance reference. Three desiderata guide interpretation: serving architecture readiness, memory discipline, and multi-node scaling. On Qwen3-0.6B, vllm-metal alone more than doubles throughput on both workloads from concurrency 1 to 16. CUDA vLLM and SGLang show stronger concurrency scaling on the same prompts. Explicit memory budgets do not guarantee memory headroom: two stacks complete every request while memory use approaches physical capacity and throughput declines. The newer model architectures have narrower engine support. Their evaluated implementations match the fidelity reference. Only three stacks satisfy the completion, fidelity, and model-coverage gates. Comparisons on larger dense and MoE models reinforce the importance of scheduling prompt processing alongside ongoing generation: vllm-metal's packed prefill-decode path maintains lower first-token latency than omlx under concurrent load. In the tested two-machine configurations, tensor parallelism over Thunderbolt RDMA scales while pipeline parallelism over TCP regresses. We release benchmark code, per-run results, and maintenance journals, supported by a workflow combining bounded agent fixes with human review.

补充信息

↑