arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Elastoformer:通过弹性模型变换实现动态适应性

Elastoformer: Enabling Dynamic Adaptivity via Elastic Model Transformation

Sudaksh Kalra, Dolly Sapra

arXiv 2609.10018首次发表:更新:

发表机构

University of Amsterdam(阿姆斯特丹大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对边缘AI动态运行条件,提出Elastoformer框架,将传统神经网络转换为弹性网络,实现运行时动态切换操作模式,显著降低计算量、延迟和内存开销,并适用于多种架构。

AI 中文摘要

边缘AI系统越来越多地采用计算机视觉应用,以实现实时的智能设备端决策。然而,这些部署面临高度动态的运行条件,延迟、功耗和内存资源的约束不断波动。深度神经网络(DNN)遵循固定的计算执行流程,缺乏适应这种变化的能力,导致在边缘场景中性能低下且效率不佳。这凸显了对不仅高效而且能在运行时动态扩展的架构的需求。在本文中,我们提出了Elastoformer:一个将传统神经网络(NN)转换为能够进行实时弹性推理的弹性NN的框架。与传统的模型包方法(需要为不同运行条件维护多个独立模型)不同,Elastoformer提供了一个单一的模块化解决方案,在运行时动态切换多种操作模式,高效适应边缘设备不断变化的计算预算,而无需管理单独模型的额外开销。实验表明,我们的框架实现了高达85%的计算FLOPs减少、50%的延迟降低和76%的内存开销减少,同时展示了该框架在Vision Transformers和CNN上的架构无关性。我们的代码可在以下网址获取:https URL。

英文摘要

EdgeAI systems are increasingly employing computer vision applications to enable intelligent, on-device decision-making in real-time. However, these deployments face highly dynamic operational conditions, with fluctuating constraints on latency, power availability, and memory resources. Deep Neural Networks (DNN), which follow fixed computational execution flows, lack the flexibility to adapt to such variability, resulting in inefficient and suboptimal performance in edge scenarios. This underscores the need for architectures that are not only efficient but also dynamically scalable at runtime. In this paper, we propose Elastoformer: A framework that transforms conventional neural networks (NN) into Elastic NN capable of real-time elastic inference. Unlike the conventional bag-of-models approach, which requires maintaining multiple independent models for different operating conditions, Elastoformer offers a single, modular solution that dynamically switches between multiple modes of operation at runtime, adapting efficiently to the changing computational budgets of edge devices without the overhead of managing separate models. Experiments reveal that our framework achieves up to 85% reduction in computation FLOPs, 50% reduction in latency and 76% reduction in memory overhead, while showcasing the architecture agnostic nature of the framework across both Vision Transformers and CNNs. Our code is available at https://github.com/sudaksh14/Elastoformer.

CommentsPublished at SEC'25

Journal refSEC 2025: Proceedings of the Tenth ACM/IEEE Symposium on Edge Computing Article No.: 10, Pages 1 - 14

DOI:10.1145/3769102.3770612

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑