WARD:面向可靠边缘AI的运行时工作负载自适应视觉Transformer框架
WARD: Runtime Workload-Adaptive Vision TRansformer Framework for Dependable Edge AI
- Humboldt University of Berlin(柏林洪堡大学)
- Tallinn University of Technology(塔林理工大学)
- Brandenburgische Technische Universität Cottbus-Senftenberg(科特布斯-森夫滕贝格勃兰登堡工业大学)
- Shahid Bahonar University of Kerman(克尔曼沙希德巴霍纳尔大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
WARD提出运行时自适应视觉Transformer框架,通过子网络划分、可靠性感知持续学习和动态模式调度,在FPGA加速器上实现高容错(故障率1.79%)和低开销(面积<5%)。
中文摘要 AI 辅助
边缘部署的AI系统在动态变化的功耗预算、可靠性要求和输入分布下运行,需要持续适应。这些条件出现在长期运行的边缘AI应用中,包括自主系统、工业监控和卫星机载智能。现有的容错方法假设静态运行条件,而持续学习技术在在线适应过程中忽视了并发的硬件故障。此外,运行时自适应可靠性框架在可编程AI加速器上的实际部署仍未得到充分探索。本文提出WARD,一个运行时自适应视觉Transformer框架,结合通道级子网络划分、可靠性感知的持续学习和动态运行模式调度,以根据运行时条件联合优化性能、容错能力和适应性。两个物理隔离的子网络在四种运行模式(即全精度模式、低功耗模式、高可靠性模式和自适应模式)下执行,动态调整计算成本和可靠性,同时确保满足实时要求的不间断推理。为验证所提框架的实际可部署性,WARD在基于FPGA的轻量级加速器上实现,并扩展了用于模式调度和资源管理的运行时硬件支持。实验结果表明,所提出的分离架构在高误码率下实现了仅1.79%的网络级故障率。硬件实现仅产生不到5%的面积开销,并支持在几个时钟周期内完成运行时模式转换,证明自适应可靠性管理可以以可忽略的实现开销集成到可编程边缘AI加速器中。
英文摘要
Edge-deployed AI operate under dynamically changing power budgets, reliability requirements, and input distributions, requiring continuous adaptation. Such conditions arise in long-running edge AI applications, including autonomous systems, industrial monitoring, and satellite onboard intelligence. Existing fault-tolerant methods assume static operating conditions, whereas continual learning techniques neglect concurrent hardware faults during online adaptation. Moreover, the practical deployment of runtime-adaptive reliability frameworks on programmable AI accelerators remains largely unexplored. This paper presents WARD, a runtime-adaptive Vision Transformer framework that combines channel-wise subnetwork partitioning, reliability-aware continual learning, and dynamic operating-mode scheduling to jointly optimize performance, fault tolerance, and adaptation according to runtime conditions. Two physically isolated subnetworks execute under four operating modes (i.e. Full-Precision Mode, Low-Power Mode, High-Reliability Mode, and Adaptive Mode) that dynamically adjust computational cost and reliability while ensuring uninterrupted inference for real-time requirements. To validate the practical deployability of the proposed framework, WARD is implemented on a lightweight FPGA-based accelerator extended with runtime hardware support for mode scheduling and resource management. Experimental results demonstrate that the proposed split architecture achieves a network-level failure rate of only 1.79% under high Bit Error Rates. The hardware implementation incurs less than 5% area overhead and supports runtime mode transitions within few clock cycles, demonstrating that adaptive reliability management can be integrated into programmable edge AI accelerators with negligible implementation overhead.