arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

YOLOv14:具备自适应多视图表示的统一跨域实时目标检测

YOLOv14: Adaptive Real-Time Object Detection for Diverse Imaging Conditions

Jian Lu, Jinling Jia, Jone Yawl, Chenbin Zhang

arXiv 2608.04720首次发表:更新:

发表机构

Nanjing University of Posts and Telecommunications(南京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

YOLOv14是具备自适应多视图表示的统一跨域实时目标检测框架,通过四项创新解决非理想输入下的性能衰减问题,在COCO及多类特殊场景基准上实现精度与速度的显著提升。

AI 中文摘要

实时目标检测器在受控条件下可达到出色的准确率,但在非理想输入(鱼眼畸变、游戏渲染角色、航拍视角、360°全景)上性能会急剧下降。本文提出YOLOv14,这一统一检测框架通过四项协同创新应对上述挑战:(1)可变形区域注意力(D-AAttn)将刚性注意力网格替换为学习得到的二维变形场,使模型能在几何畸变下自适应采样;(2)游戏到真实域自适应(Game2Real Domain Adaptation)通过自适应实例归一化(AdaIN)和对抗域混淆对齐渲染游戏与摄影图像的特征分布,实现游戏角色可被检测为真实人类;(3)多视图条件注入学习得到的视角嵌入至骨干网络,搭配跨视图对比损失拉近不同视角下同一类别的特征;(4)自适应增强策略自动对每个输入的场景类型分类并匹配最优增强方式,动态尺度路由(DynamicScaleRouter)学习每个输入的特征金字塔权重。YOLOv14在COCO val2017数据集上以2.91毫秒(T4 GPU)的推理速度达到49.1 mAP,在鱼眼、全景、无人机、游戏角色基准上分别实现+4.1 mAP、+6.6 mAP、+6.4 mAP、+26.1 mAP的显著提升。

英文摘要

Real-time object detectors achieve remarkable accuracy under controlled conditions, yet degrade sharply on non-ideal inputs-fisheye distortion, game-rendered content, aerial views, and 360°panoramas. We present YOLOv14, a unified adaptive detection framework that addresses these variations through four complementary mechanisms, formalized under a novel Adaptive Routing and Modulation (ARM) paradigm. Unlike conventional unsupervised domain adaptation, our approach employs Target-Prior Guided Source-Domain Augmentation(TP-SDA), using only 50 unlabeled target images offline to estimate style statistics, while adversarial alignment serves as a lightweight regularizer rather than the primary adaptation driver. Together, these components enable YOLOv14 to achieve 49.1 mAP on COCO val2017 at 2.91 ms (T4 GPU), with substantial gains of +4.1 (fisheye), +6.6 (panorama), +6.4 (drone), and +26.1 (gamestylized) mAP over YOLOv12s. Crucially, we validate generalization on real-world game screenshots (GTA-V, Unity), achieving +14.2 mAP, confirming practical transferability beyond synthetic benchmarks. Code and models are released at https://github.com/zhangcbb/yolov14.

CommentsSorry, we need to evaluate and revise the paper more scientifically

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑