arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02689cs.CV

GaLe:内存高效的全局近似与局部精确特征

GaLe: memory-efficient Global Approximate and Local Exact features

  • Fondazione Bruno Kessler (FBK)(布鲁诺·凯塞勒基金会)

机构由 AI 辅助整理,请以论文原文为准。

Alberto Ancilotto, Elisabetta Farella

AI总结:

该研究提出GaLe技术,通过划分特征图为局部精确与全局近似表示,在不重训预训练网络的情况下,实现受限设备上的高效部署,在ImageNet上匹配精确推理性能且显著降低内存、提升速度,适用于多类计算机视觉任务。

AI中文摘要:

嵌入式设备通常缺乏配备GPU的机器的资源,而现有的推理方法要么存在高计算开销(基于补丁的方法),要么存在精度损失(基于近似的方法)。我们提出GaLe,一种内存高效的技术,可在不重新训练的情况下将预训练网络部署在受限设备上。GaLe将特征图划分为两个部分:保留精细细节的局部精确(Local exact, Le)表示,以及保留长程依赖的全局近似(Global approximate, Ga)表示。与标准分块不同,GaLe支持混合CNN-Transformer模型中常见的全局操作和注意力机制。在ImageNet上验证,我们的方法达到了精确推理的性能,与基于补丁的推理相比,在Cortex-M33上实现了高达65%的加速和90%的RAM减少。我们进一步证明GaLe在分类、检测和生成任务中的通用性,突出其作为资源高效架构设计基础的潜力。

英文摘要:

Embedded devices typically lack the resources of GPU-equipped machines, and existing inference methods suffer from either high computational overhead (patch-based) or accuracy loss (approximation-based). We propose GaLe, a memory-efficient technique that enables the deployment of pretrained networks on constrained devices without retraining. GaLe partitions feature maps into two components: a local exact (Le) representation that preserves fine details and a global approximate (Ga) representation that retains long-range dependencies. Unlike standard tiling, GaLe supports global operations and attention mechanisms found in hybrid CNN-transformer models. Validated on ImageNet, our method matches exact-inference performance while achieving up to 65% speedup and 90% RAM reduction on a Cortex-M33 compared to patch-based inference. We further demonstrate GaLe's versatility across classification, detection, and generation tasks, highlighting its potential as a foundation for resource-efficient architecture design.

↑