arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向嵌入式系统的流学习:流学习方法内存消耗的基准测试

Towards Stream Learning on Embedded Systems: Benchmarking the Memory Consumption of Stream Learning Methods

Sebastian Buschjäger, Nuwan Gunasekara, Heitor Murilo Gomes

arXiv 2608.30923首次发表:更新:

发表机构

Lamarr Institute; Halmstad University; Victoria University of Wellington(拉马尔研究所; 哈尔姆斯塔德大学; 惠灵顿维多利亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对嵌入式系统资源稀缺问题,通过6463次实验基准测试7种流分类器,发现自适应集成与增量树的内存故障模式,呼吁将有界资源使用设为流学习设计目标并提出相关API。

AI 中文摘要

流学习通常通过预测性能和对概念漂移的适应性进行评估。然而,流学习器的持续运行还需要在长数据流上具备可预测且有界的资源使用,当学习从服务器迁移到传感器附近的嵌入式系统(内存和处理资源稀缺)时,这一要求变得尤为关键。但在最先进的流学习中,研究重点往往集中在概念漂移适应上,而资源使用通常只是评估的附带结果。为缩小这一差距,我们在128 KiB至约8 MiB的模型大小预算下,对7种代表性流分类器在13个真实和合成数据流上进行了基准测试,共开展6463次实验。我们测量了故障感知准确率、峰值模型大小、预算耗尽时间以及预测加更新延迟。结果揭示了两种不同的资源故障模式:自适应集成模型会因初始内存占用几乎立即超过小预算,即便之后大小保持稳定;增量树可适配初始预算,但会在长数据流中持续增长,Hoeffding树(HT)和极快决策树(EFDT)的中位数增长倍数分别为7.37和5.87。显式压缩方法是最小预算下唯一可行的选择,但当预算更大使自适应集成模型具备竞争力时,通常会被后者超越。因此,许多最先进的方法仅部分适用于嵌入式系统或长运行系统。我们呼吁流学习社区将有界资源使用与漂移适应一同作为一流设计目标,并提出了实现该目标的具体步骤,包括一个流学习器可显式暴露并遵循资源预算的API。

英文摘要

Stream learning is commonly evaluated through predictive performance and adaptation to concept drift. However, sustained operation of a stream learner also requires predictable and bounded resource usage even on long streams. This requirement becomes even more critical when learning moves from servers to near-sensor embedded systems where memory and processing are scarce resources. In state-of-the-art stream learning, however, we perceive a strong focus on concept drift adaptation, whereas resource usage is often an evaluation byproduct. To close this gap, we benchmark seven representative stream classifiers on 13 real and synthetic streams under model-size budgets from 128\,KiB to approximately 8\,MiB. Our benchmark comprises a total of 6,463 experiments. We measure failure-aware accuracy, peak model size, time to budget exhaustion, and prediction-plus-update latency. The results reveal two distinct resource failure modes. Adaptive ensembles can exceed small budgets almost immediately because of their initial footprint, even when their size remains stable thereafter. Incremental trees can fit initially but grow throughout a long stream, with HoeffdingTrees (HT) and Extremely Fast Decision Trees (EFDT) increasing by median factors of 7.37 and 5.87. Explicitly compact methods remain the only viable option under the smallest budgets, but are usually overtaken as larger budgets make adaptive ensembles competitive. Hence, many state-of-the-art methods are only partially applicable in embedded systems or for long-running systems. We therefore call on the stream-learning community to make bounded resource usage a first-class design objective alongside drift adaptation, and propose concrete steps toward this goal, including an API through which stream learners can explicitly expose and respect resource budgets.

Comments7 pages double-column + appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑