arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FORGE:面向微控制器上纯整数视觉模型的仅前向测试时自适应方法

FORGE: Forward-Only Test-Time Adaptation for Integer-Only Vision Models on Microcontrollers

Muhammad Rehan, Haider Ali, Muhammad Ali Munir, Moaz Amjad

arXiv 2609.01683首次发表:更新:

发表机构

SEECS, NUST; FAST-NUCES(国家科技大学SEECS学院; FAST-NUCES大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FORGE是一种可在微控制器纯整数视觉模型上运行的仅前向测试时自适应方法,通过重新归一化融合卷积的输出恢复自适应能力,能耗低且泛化性好。

AI 中文摘要

部署在微控制器(MCU)上的视觉模型被量化为纯整数算术,且在仅推理的运行时环境中运行,该环境不具备反向传播所需的机制,而反向传播是使模型适应其在现场遇到的分布偏移(传感器噪声、模糊、光照等)的标准工具。现有的仅前向测试时自适应(TTA)方法要么仅在服务器或边缘GPU级别的模型上运行(并非真正的微控制器整数执行),要么需要整数部署中已融合去除的批归一化(BN)层。我们提出了一种仅前向TTA方法,该方法可在已部署的、BN融合的纯整数卷积网络上运行。关键发现是,将BN融合到前序卷积中(这是整数推理的必要步骤)会破坏基于归一化的自适应所依赖的统计量。我们通过仅使用前向传播估计值,将每个融合卷积的逐通道输出重新归一化至其干净训练统计量来恢复自适应能力。该方法:(i)恢复了基于梯度的TENT的大部分精度提升(+20.9个百分点对比+24.9个百分点),且与仅前向BN自适应方法的表现相当,同时是唯一能在融合的纯整数模型上运行的方法;(ii)仅需对21层中的3层进行自适应(在未看到测试损坏的情况下选择)即可恢复93%的收益;(iii)可通过批大小缩放的动量在单样本流式场景下运行;(iv)可跨三个数据集(最多200个类别)和两种架构泛化。我们验证了比特精确的int8卷积执行,并将其部署在ESP32-S3上,使用Nordic PPK2功率分析器测量,仅前向自适应(围绕int8卷积的轻量级fp32重新校准)仅消耗8.3 mJ(占推理能耗的6.8%),在已部署的SIMD优化模型上耗时21.9 ms:仅前向自适应在真实微控制器上的成本极低。

英文摘要

Vision models deployed on microcontrollers (MCUs) are quantized to integer-only arithmetic and run in inference-only runtimes that do not carry the machinery backpropagation needs: the standard tool for adapting a model to the distribution shift (sensor noise, blur, lighting) it meets in the field. Existing forward-only test-time adaptation (TTA) methods either run only on server- or edge-GPU-class models (not true microcontroller integer execution), or require the batch-normalization (BN) layers that integer deployment fuses away. We present a forward-only TTA method that operates on deployed, BN-folded, integer-only convolutional networks. The key observation is that fusing BN into the preceding convolution, a mandatory step for integer inference, destroys the statistics that normalization-based adaptation relies on. We restore adaptation by re-normalizing each folded convolution's per-channel output to its clean training statistics, using only forward-pass estimates. The method (i) recovers most of gradient-based TENT's accuracy gain (+20.9 vs. +24.9 points) and matches forward-only BN adaptation, while being the only method that runs on a folded integer-only model; (ii) needs to adapt only 3 of 21 layers (selected without seeing the test corruptions) to recover 93% of the benefit; (iii) survives single-sample streaming with a batch-size-scaled momentum; and (iv) generalizes across three datasets (up to 200 classes) and two architectures. We validate bit-exact int8 convolution execution and deploy on an ESP32-S3, where, measured with a Nordic PPK2 power profiler, the forward-only adaptation (a lightweight fp32 recalibration around the int8 convolutions) costs only 8.3 mJ (6.8% of inference energy) and 21.9 ms on the deployed SIMD-optimized model: forward-only adaptation is cheap on a real microcontroller.

Comments16 pages, 5 figures, 9 tables. Published in Transactions on Machine Learning Research (2026). OpenReview: https://openreview.net/forum?id=A45I5p25dd. Code and checkpoints: https://github.com/Rehan000/forge-tta

Journal refTransactions on Machine Learning Research, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑