arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16157cs.DC

FreeToken:面向边缘原生MoE服务的带宽自适应高效执行方案

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

Shuo Yang, Xiaoze Fan, Melissa Pan, Haocheng Xi, Zhe Wang, Shanlin Sun, Kurt Keutzer, Song Han, Matei Zaharia, Chenfeng Xu, Ion Stoica

首次发表
浏览论文内容

中文总结 AI 辅助

FreeToken是边缘原生MoE服务系统,通过协同设计服务栈实现带宽自适应执行,将个人计算机转化为弹性推理平台,可在边缘硬件上部署大参数MoE模型,拓展了本地AI服务能力。

中文摘要 AI 辅助

前沿开源权重模型日益普及,但其服务部署仍主要依赖数据中心基础设施。我们提出FreeToken,一种边缘原生MoE服务系统,该系统将个人计算机视为统一的弹性推理平台,而非小型GPU。FreeToken围绕本地AI的两大现实场景,对完整服务栈进行协同设计,包括模型布局与加载、专家驻留、CPU-GPU执行、智能体状态复用及运行时内存管理:一是智能体工作负载的执行模式持续变化,二是边缘硬件呈现异构资源特性,且各机器间的资源平衡存在差异。FreeToken不采用固定卸载策略,而是持续将计算与模型状态映射到实际可用资源。它支持20余种MoE模型,以及从8GB笔记本电脑GPU到单工作站GPU等硬件上的真实编码与工具使用智能体;更重要的是,它拓展了这些机器的实际服务能力,从笔记本电脑上的35B模型,到游戏台式机上的284B模型,再到单工作站GPU上的753B GLM-5.2模型。FreeToken将开源权重转化为可部署的本地软件,使用户现有机器成为承载前沿规模智能的实用平台。我们在该http URL发布此系统。

英文摘要

Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack, including model layout and loading, expert residency, CPU--GPU execution, agentic state reuse, and runtime memory management, around two realities of local AI: agent workloads continuously change their execution pattern, and edge hardware exposes heterogeneous resources whose balance differs from machine to machine. Rather than committing to a fixed offloading strategy, FreeToken continuously maps computation and model state onto the resources actually available. FreeToken supports more than 20 MoE models and real coding and tool-using agents across hardware ranging from an 8GB laptop GPU to a single workstation GPU. More importantly, it changes what these machines can practically serve, from a 35B model on a laptop to a 284B model on a gaming desktop and the 753B GLM-5.2 on a single workstation GPU. FreeToken turns open weights into deployable local software, making the machines users already own a practical platform for frontier-scale intelligence. We release the system at flashml.ai.

↑