arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03149cs.ETcs.ARcs.LG

RACE-AIMC:面向边缘端异构模拟内存内加速器的选择性推理

RACE-AIMC: Selective Inference for Heterogeneous Analog In-Memory Accelerators at the Edge

Osama Yousuf, Martin Lueker-Boden

首次发表
浏览论文内容

中文总结 AI 辅助

RACE-AIMC框架为边缘端异构模拟内存内加速器,通过离线选最优加速器并计算错误上界、在线轻量检查的方式,在保证准确率的同时降低了能耗。

中文摘要 AI 辅助

模拟内存内计算(AIMC)通过直接在存储阵列内部执行算术运算来加速神经网络推理,无需在内存与处理器之间来回传输权重,从而节省了能耗。但存储权重的物理器件存在缺陷:编程误差、电噪声、有限分辨率的转换器以及完全损坏的单元都会扭曲计算,且每个物理芯片的扭曲方式各不相同。拥有多个此类芯片的设计者面临两难选择:运行所有芯片并融合结果(安全但能耗浪费),或盲目信任单个芯片(成本低但无法保证错误频率)。本文提出RACE-AIMC(面向AIMC的风险感知认证集成),这一框架通过统计方法而非猜测解决该选择问题。在离线阶段,RACE-AIMC研究一组物理加速器,为给定能耗预算选择最优的单个加速器,并计算该加速器在给出答案时错误频率的数学精确上界。在线阶段,仅激活该单个加速器,通过轻量检查决定是否接受其答案或转向备选方案。在使用有噪声权重映射和多次独立测试运行的模拟中,所有认证边界均保持在10%的错误目标以内(平均边界为7.83% ± 0.89%,70.88% ± 0.98%的输入被直接回答)。所得系统达到了干净数字基线的准确率,同时相对于始终运行集合中所有加速器的情况,建模能耗降低了69.02%。

英文摘要

Analog in-memory computing (AIMC) speeds up neural-network inference by doing the arithmetic directly inside a memory array, instead of shuttling weights back and forth between memory and a processor. This saves energy, but the physical devices that store the weights are imperfect: programming errors, electrical noise, limited-resolution converters, and outright broken cells all distort the computation, and every physical chip is distorted in its own way. A designer with several such chips available faces an uncomfortable choice: run all of them and combine the answers (safe, but wasteful of energy), or trust a single chip blindly (cheap, but with no guarantee on how often it is wrong). This paper introduces RACE-AIMC (Risk-Aware Certified Ensemble for AIMC), a framework that resolves this choice with statistics rather than guesswork. Offline, RACE-AIMC studies a pool of physical accelerators, picks the single best one for a given energy budget, and computes a mathematically exact upper bound on how often that accelerator will be wrong when it chooses to answer. Online, only that one accelerator is switched on; a lightweight check decides whether to accept its answer or defer to a fallback. In our simulations using a noisy weight mapping and multiple independent test runs, every certified bound stayed under a 10% error target (mean bound 7.83% +- 0.89%, with 70.88% +- 0.98% of inputs answered directly). The resulting system matches the accuracy of a clean digital baseline while cutting modeled energy use by 69.02% relative to always running every accelerator in the pool.

发表机构

  • WD Research(西部数据研究机构)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑