发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出阶梯式专家混合框架,结合弹性结构与稀疏门控,实现模型同时适应部署约束和任务需求,支持1-4B参数灵活配置,在知识密集型基准上比密集模型准确2-5%,且延迟相当,节省存储并提升服务灵活性。
AI 中文摘要
训练大型语言模型(LLM)资源密集,且使其适应具有不同计算约束的多种部署场景仍具挑战性。虽然弹性架构能够实现灵活的模型部署,稀疏激活模型允许输入自适应计算,但现有方法将这些维度独立处理。此外,面向设备端边缘推理的模型需要符合服务设备的存储和计算限制。在本文中,我们引入了一个统一框架,将弹性结构与稀疏门控架构相结合,以创建同时适应部署约束和任务需求的模型。我们的方法采用一个模型主干,其同时基于上下文和目标效率规格进行条件化,从而在推理时实现对精度-效率权衡的细粒度控制。该模型学习在弹性嵌套的子网络内激活任务相关参数,使单个模型能够跨越多个容量点,同时保持输入自适应路由。通过实验,我们证明可以创建一个模型,该模型允许灵活使用10亿、20亿、30亿、40亿参数,同时比其密集对应模型更准确(在知识密集型基准上提高2-5%),并且与其静态版本相当,同时提供与密集模型相似的延迟指标。总体而言,我们通过共享模型参数节省设备磁盘空间,允许基于DRAM和可用计算进行灵活服务,同时提供更准确的结果。
英文摘要
Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches treat these dimensions independently. Moreover, models catered towards on-device edge inference need to conform to the memory and compute limitations of the serving devices. In this paper, we introduce a unified framework that combines elastic structures with sparsely gated architectures to create models that adapt simultaneously to both deployment constraints and task requirements. Our approach employs a model backbone that conditions on both the context and target efficiency specifications, enabling fine-grained control over the accuracy-efficiency trade-off at inference time. The model learns to activate task-relevant parameters within elastically-nested sub-networks, allowing a single model to span multiple capacity points while maintaining input-adaptive routing. Through experiments we demonstrate that we can create a model that allows the flexibility to use 1,2,3,4 billion parameters while being more accurate than their dense counter-parts (2-5\% on knowledge-intensive benchmarks) and at par with their static versions while delivering similar latency metrics as dense models. Overall, we save on device disk space by sharing the model parameters, allow flexibility of serving based on DRAM and compute available while delivering more accurate results.
CommentsApple Foundation Models, 15 Pages, Edge LLMs