发表机构
Architect Labs(建筑师实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出端到端AI系统实现软硬件协同设计验证,2周内自主生成Redwood AI加速器,其性能功耗比优于Jetson,是首个AI端到端设计的可生产AI加速器。
AI 中文摘要
现代AI工作负载及其运行硬件的演化 timescale 不同:架构定义先于量产芯片数年,而目标工作负载数月内就会变化。因此,设计决策在深度不确定性下做出,需付出双重代价:一是为对冲风险增加通用性,二是当新工作负载难以适配冻结的芯片时付出代价。随着摩尔定律停滞,专业化是性能功耗比的主要来源,需要与工作负载节奏匹配的设计周期。我们提出一种端到端AI系统,将软件到芯片的堆栈压缩为单一优化循环,硬件与软件在同一目标下协同设计与验证。首个演示是Redwood,一款为物理AI设计的单批次、低功耗、超低延迟推理前沿AI加速器。由两名人类架构师给出高级规范后,该系统在无人工干预的情况下,在不到2周内自主生成性能模型、RTL设计、UVM环境、形式化证明、固件和内核。每个模块通过商用EDA工具、我们的专有形式化引擎和硬件在环验证达到95%覆盖率。规范变更在48小时内重新验证并部署到硬件。其超低功耗FPGA变体Redwood Nano可运行Llama和Qwen等数十亿参数模型。投射到三星8nm(Jetson Orin Nano的工艺等级)时,Redwood在相同模型上相比实测Jetson基线,吞吐量达1.75倍,功耗降低1.9倍,性能功耗比提升3.4倍。在Redwood上运行的Qwen还助力设计下一代Redwood,这是递归自我改进的早期步骤。据我们所知,这是首个由AI系统端到端设计并运行现代AI模型的可生产AI加速器。
英文摘要
Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by years, while target workloads shift in months. Design decisions are therefore committed under deep uncertainty and paid for twice, once in the generality added as a hedge, and again when new workloads map poorly onto frozen silicon. As Moore's Law stagnates, specialization is the main remaining source of performance-per-watt and demands a design cycle that runs at the cadence of the workloads. We present an end-to-end AI system that collapses the software-to-silicon stack into a single optimization loop, where hardware and software are co-designed and verified under one objective. Its first demonstration is Redwood, a frontier AI accelerator built for single-batch, low-power, ultra-low-latency inference for physical AI. From a high-level specification by two human architects, the system autonomously generated the performance model, RTL design, UVM environments, formal proofs, firmware, and kernels in under two weeks with no human intervention below the specification. Every block reached 95% coverage via commercial EDA tools, our proprietary formal engine, and hardware-in-the-loop validation. Specification changes were reverified and redeployed to hardware in under 48 hours. Redwood Nano, its ultra-low-power FPGA variant, runs multi-billion-parameter models like Llama and Qwen. Projected onto Samsung 8 nm, the Jetson Orin Nano's process class, Redwood delivers 1.75x the throughput at 1.9x lower power, a 3.4x performance-per-watt gain against a measured Jetson baseline on the same models. Qwen running on Redwood also helped design next-generation Redwood, an early step toward recursive self-improvement. To our knowledge, this is the first production-worthy AI accelerator designed end-to-end by an AI system and running a modern AI model.
Comments7 Pages, and 15 figures