发表机构
CEA-Leti, Mines Saint-Etienne, Equipe Commune SAS; Univ. Grenoble Alpes, CEA-Leti; Mines Saint-Etienne, CEA-Leti, Centre CMP, Equipe commune SAS(CEA-莱蒂、圣艾蒂安 Mines、共同团队 SAS; 格勒诺布尔阿尔卑斯大学、CEA-莱蒂; 圣艾蒂安 Mines、CEA-莱蒂、CMP 中心、共同团队 SAS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在RISC-V单核上进行Float16设备端训练,利用Zfh和Zvfh扩展提出开源框架,能减内存占用、促迁移学习,基于AIfES构建,强调Zfh低开销并讨论Zvfh架构。
AI 中文摘要
利用标准的RISC-V扩展(Zfh(标量float16)和Zvfh(向量float16)),本文提出了一个开源框架,以在资源受限的RISC-V单核上实现完整的设备端训练。与使用float32相比,该方法可使内存占用减少约50%,且模型性能下降最小。通过整合层冻结功能,还促进了迁移学习和微调场景。基于AIfES构建,强调了Zfh在RV64GC超标量乱序FPGA软核上的低面积开销。最后讨论了在同一RISC-V内核中Zvfh实现的架构。
英文摘要
By leveraging standard RISC-V extensions, namely Zfh (scalar float16) and Zvfh (vector float16), this work proposes an open-source framework to enable complete on-device training on resource-constrained RISC-V single-core. Our approach allows memory footprint reduction by about 50% as compared to using float32 and with minimal model performance degradation. We also facilitate transfer learning and fine-tuning scenarios by incorporating layer-freezing capabilities. Our work builds onto AIfES, an open-source, modular and generic DNN training and inference framework for embedded systems that can be extended with custom hardware-specific functions. The benefits of float16 is further emphasized by outlining the low area overhead of Zfh on a RV64GC super-scalar out-of-order FPGA softcore (+1.15% LUT6 and +0.05% FF at 175MHz). Finally, we discuss the architecture of a Zvfh implementation within the same RISC-V core.
CommentsAccepted at IEEE PRIME 2026