arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02608cs.DC

将智能与推理分离:边缘原生AI计算的标准

Separating Intelligence from Inference: A Standard for Edge-Native AI Computing

Venkat Vinjam, Krishnaiah Narukulla

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出将AI的训练(智能)与推理分离的原则,设计了两类边缘设备及相关架构栈,可大幅降低大型语言模型推理的能源消耗与碳排放。

中文摘要 AI 辅助

人工智能行业已构建了价值3000亿美元的集中式数据中心基础设施,用于服务大型语言模型推理这类工作负载,而该工作负载在架构上并不要求集中化。本文阐明了当代AI基础设施的核心架构低效性:将模型训练(本质上是集中式、资本密集型、每个模型版本仅需一次)与模型推理(可并行化、对延迟敏感、每次查询需重复执行)在同一物理硬件上混为一谈。我们提出分离原则:智能在集中环境中训练并以软件形式交付;推理则在靠近数据源的边缘硬件上执行。我们量化了文明尺度下的能源影响,结果显示,与当前的集中式实践相比,面向每日10亿用户的全边缘推理架构每年可节省约19太瓦时(TWh)能源和7.3兆吨二氧化碳(CO₂)。我们明确了两类新设备:个人AI计算机(PAC)和企业AI工作站(CAW),并给出了具体的硬件层级、内存带宽要求、热设计功耗范围及软件接口。随后,我们描述了包含八个组件的参考架构栈,这些组件涉及权重分布、主权感知路由、热自适应量化、多租户资源管理、联邦网络推理、加密溯源、隐私保护遥测以及分布式上下文窗口扩展。其中若干组件是第一作者正在申请的美国专利主题,在此作为候选开放架构原则呈现。

英文摘要

The artificial intelligence industry has constructed a USD 300 billion centralized data center infrastructure to serve a workload, large language model inference, that does not architecturally require centralization. This paper articulates the central architectural inefficiency of contemporary AI infrastructure: the conflation of model training (irreducibly centralized, capital-intensive, one-time per model version) with model inference (parallelizable, latency-sensitive, recurring per query) on the same physical hardware. We propose the separation principle: intelligence is trained centrally and shipped as software; inference executes on hardware near the data source, at the edge. We quantify the energy implications at civilizational scale and show that a fully edge-resident inference architecture for one billion daily users saves approximately 19 TWh per year and 7.3 megatons of CO2 annually relative to current centralized practice. We specify two new device classes, the Personal AI Computer (PAC) and the Corporate AI Workstation (CAW), with concrete hardware tiers, memory bandwidth requirements, thermal envelopes, and software interfaces. We then describe a reference architectural stack of eight components addressing weight distribution, sovereignty-aware routing, thermal-adaptive quantization, multi-tenant resource management, federated network inference, cryptographic provenance, privacy-preserving telemetry, and distributed context window extension. Several components are the subject of pending United States patent applications by the first author and are presented here as candidate open architectural principles

补充信息

↑