JEDI:面向卫星影像高效农田分割的JEPA到边缘蒸馏
JEDI: JEPA-to-Edge Distillation for Efficient Cropland Segmentation from Satellite Imagery
浏览论文内容
中文总结 AI 辅助
JEDI提出两阶段蒸馏框架,将大型I-JEPA教师表示迁移至紧凑SegFormer学生,通过持续特征对齐在CalCROP21上以4.04M参数达68.0 mIoU,接近教师性能并优于多种蒸馏基线。
中文摘要 AI 辅助
大型视觉模型为遥感分割提供了有用的表示,但往往因成本过高而难以在卫星或田间边缘部署。现有的特征级蒸馏方法也倾向于假设教师和学生架构相似,并且通常在任务训练开始时停止特征对齐。我们提出了JEDI(JEPA到边缘蒸馏),一个两阶段框架,将表示从大型I-JEPA Vision Transformer教师转移到紧凑的SegFormer学生。首先,JEDI通过跨架构投影和空间对齐将学生的终端表示与教师的令牌空间对齐。然后,在整个任务适应过程中,联合优化监督分割、温度缩放响应蒸馏和持续特征对齐。在CalCROP21上,JEDI-B0以4.04M参数实现了68.0的平均交并比(mIoU),比独立学生提高了16.0个百分点,并接近由639M参数教师实现的70.0 mIoU的2.0个百分点以内。我们分别评估了具有4.04M、14.33M和28M参数的SegFormer B0、B1和B2学生。在所有三种变体中,JEDI在相同教师-学生设置下始终优于基于响应、结构、通道和关系的蒸馏基线。这些结果表明,持续表示对齐在激进压缩下尤其有价值,显著减少模型大小和计算量,同时保持分割性能。
英文摘要
Large vision models provide useful representations for remote-sensing segmentation but are often too expensive for deployment at the satellite or field edge. Existing feature-level distillation methods also tend to assume similar teacher and student architectures and often stop feature alignment when task training begins. We introduce JEDI (JEPA-to-Edge Distillation), a two-stage framework that transfers representations from a large I-JEPA Vision Transformer teacher to a compact SegFormer student. First, JEDI aligns the student's terminal representation with the teacher's token space using cross-architecture projection and spatial alignment. It then jointly optimizes supervised segmentation, temperature-scaled response distillation, and persistent feature alignment throughout task adaptation. On CalCROP21, JEDI-B0 achieves 68.0 mean Intersection-over-Union (mIoU) with 4.04M parameters, improving over the standalone student by 16.0 points and coming within 2.0 points of the 70.0 mIoU achieved by the 639M-parameter teacher. We evaluate SegFormer B0, B1, and B2 students with 4.04M, 14.33M, and 28M parameters, respectively. Across all three variants, JEDI consistently outperforms response-, structure-, channel-, and relational-distillation baselines under the same teacher-student setting. These results show that persistent representation alignment is especially valuable under aggressive compression, substantially reducing model size and computation while preserving segmentation performance.
发表机构
- University of California, Riverside(加州大学河滨分校)
机构由 AI 辅助整理,请以论文原文为准。