arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20441cs.CVcs.LG

跨架构基础模型蒸馏用于边缘洪水分割

Cross-Architecture Foundation-Model Distillation for Edge Flood Segmentation

Fabian Schmalstieg, Karsten Mueller, Wojciech Samek

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过跨架构蒸馏,将3亿参数教师模型的知识迁移至70万参数学生模型,利用无标注数据扩展训练集,在保持性能的同时实现边缘设备上的高效洪水分割。

中文摘要 AI 辅助

地理空间基础模型能够提供强大的洪水分割性能,但其规模限制了在内存受限的边缘硬件上的部署。我们将一个在252个手动标注的Sen1Floods11训练场景上微调的3亿参数Prithvi-EO-2.0教师模型,蒸馏为一个70万参数的EfficientViT-B0学生模型。教师模型监督额外的无标注Sentinel-2影像,使得学生训练集无需新增手动标注即可扩大。在匹配的252个场景预算下,教师监督训练与直接训练竞争力相当,并在测试配置中提升了STURM-Flood性能;几何匹配的对照组表明,仅标签来源无法解释该差异。将教师监督池扩展到2500个场景缩小了剩余的学生-教师差距:浮点学生模型在Sen1Floods11测试集上的水交并比达到0.787,而教师为0.822;在我们的评估协议下,学生模型在STURM-Flood上与教师持平,但在WorldFloods-v2上仍低于教师。经过激活替换和量化感知训练后,学生模型作为1.5兆字节的8位整数(INT8)TensorRT引擎在Jetson Xavier NX上运行,每张512×512图像的图形处理单元(GPU)计算时间为5.57毫秒,运行时设备内存约为14兆字节。固定的修正归一化差异水指数(MNDWI)阈值在两个干净的外部基准上与两个模型竞争力相当,因此我们将这些基准视为泛化测试,而非学习模型优于光谱规则的证据。结果支持以下结论:基础模型监督可以将固定的手动标注预算放大为大幅扩大的训练集,并产生紧凑、可部署的边缘模型。

英文摘要

Geospatial foundation models can provide strong flood-segmentation performance, but their size limits deployment on memory-constrained edge hardware. We distill a 300-million-parameter Prithvi-EO-2.0 teacher, fine-tuned on the 252 manually labeled Sen1Floods11 training scenes, into a 0.7-million-parameter EfficientViT-B0 student. The teacher supervises additional unlabeled Sentinel-2 imagery, allowing the student training set to grow without new manual annotations. At the matched budget of 252 scenes, teacher-supervised training is competitive with direct training and improves STURM-Flood performance across tested configurations; a geometry-matched control shows that label source alone does not explain the difference. Scaling the teacher-supervised pool to 2,500 scenes narrows the remaining student--teacher gap: the float student reaches 0.787 water intersection over union on the Sen1Floods11 test split against 0.822 for the teacher, matches the teacher on STURM-Flood under our evaluation protocol, and remains below it on WorldFloods-v2. After activation replacement and quantization-aware training, the student runs as a 1.5-megabyte 8-bit integer (INT8) TensorRT engine on a Jetson Xavier NX at 5.57 milliseconds of graphics processing unit (GPU) compute per 512-by-512 image, with approximately 14 megabytes of runtime device memory. A fixed modified normalized difference water index (MNDWI) threshold is competitive with both models on the two clean external benchmarks, so we interpret those benchmarks as generalization tests rather than as evidence of learned-model superiority over a spectral rule. The results support the conclusion: foundation-model supervision can amplify a fixed manual annotation budget into a substantially larger training set and yield a compact, deployable edge model.

补充信息

↑