发表机构
Hunan University; Karlsruhe Institute of Technology(湖南大学; 卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FUSEye通过重叠网格视图、零初始化适配器和跨投影一致性融合,以极轻训练成本将COCO预训练YOLO26-x高效转化为鱼眼检测器,在WoodScape基准上显著提升性能。
AI 中文摘要
鱼眼相机为移动机器人提供了单传感器、低成本的周围环境视图,然而从业者常规复用的COCO预训练检测器在其上表现不佳:强烈的径向畸变扭曲了局部图像结构,而边界压缩使物体缩小至近乎不可见的大小。全微调在很大程度上弥补了这一差距,但需要大量的鱼眼标签和计算资源。我们提出了FUSEye,一种轻训练框架,将冻结骨干的COCO预训练超大YOLO26检测器(YOLO26-x)转变为鱼眼检测器。FUSEye在更新插入模块和预训练检测头的同时,仅增加约22.7万个新参数。它在三个因果关联的层面上解决迁移差距。在输入层面,重叠网格视图生成和框重映射(GridViews)扩大了压缩的边界区域。在特征层面,零初始化残差适配器(Z-Adapters)纠正了畸变引起的特征错位。在决策层面,学习的跨投影一致性融合(AgreeFusion)仅在多个视图的一致证据支持时提升低置信度检测。在WoodScape环视鱼眼基准上,FUSEye将YOLO26-x的mAP50从0.148提升至0.266,并保留了84.3%的全微调精度。此外,随机仅使用25%的标注训练图像,FUSEye达到0.2597的mAP50,保留了其全标签性能的97.6%。FUSEye还持续改进了YOLOv8-11检测器,表明该方法是架构无关的。源代码将在此https URL提供。
英文摘要
Fisheye cameras give mobile robots a single-sensor, low-cost view of their surroundings, yet the COCO-pretrained detectors that practitioners routinely reuse fail on them: strong radial distortion warps local image structure, while boundary compression shrinks objects to near-invisible sizes. Full fine-tuning closes much of the gap but requires abundant fisheye labels and compute. We present FUSEye, a training-light framework that turns a frozen-backbone COCO-pretrained extra-large YOLO26 detector (YOLO26-x) into a fisheye detector. FUSEye adds roughly 227k new parameters while updating the inserted modules and the pretrained detection head. It addresses the transfer gap at three causally linked levels. At the input level, overlapping grid view generation and box remapping (GridViews) enlarge compressed boundary regions. At the feature level, zero-initialized residual adapters (Z-Adapters) correct distortion-induced feature misalignment. At the decision level, learned cross-projection agreement fusion (AgreeFusion) promotes low-confidence detections only when they are supported by consistent evidence across multiple views. On the WoodScape surround-view fisheye benchmark, FUSEye raises YOLO26-x from 0.148 to 0.266 mAP50 and retains 84.3% fully fine-tuned accuracy. Moreover, randomly using only 25% of the labeled training images, FUSEye achieves 0.2597 mAP50, retaining 97.6% of its full-label performance. FUSEye also consistently improves YOLOv8-11 detectors, showing that the recipe is architecture-agnostic. Source code will be available at https://github.com/Su-wenya/FUSEye.
CommentsSource code will be available at https://github.com/Su-wenya/FUSEye