发表机构
Sungkyunkwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对隐私要求下训练评估流程盲态的跨城市目标检测域偏移问题,提出含类别无关目标性蒸馏与Grayworld变换的模块化流程,优化RF-DETR变体在AI City Challenge Track 6获第一。
AI 中文摘要
交通监控系统的实际部署受限于地理域偏移问题,在一个城市训练的模型应用于未见过的目标城市时性能会下降。传统域自适应方法依赖对超参数敏感的架构或直接分析目标数据,这在注重隐私的生态系统中根本无法使用,这类系统要求训练和评估流程完全盲态。在该场景下,我们探究预训练与数据增强在解决域偏移问题中的作用。具体而言,我们提出一种用于目标检测的模块化训练流程,该流程围绕两个核心正交支柱构建:(1)多数据集预训练策略,包含类别无关的目标性蒸馏,以将车辆的结构几何特征与语义分类体系解耦;(2)域鲁棒性增强流,包含一种新型Grayworld变换,该变换迫使全局注意力头剥离不稳定的色彩捷径,转而依赖鲁棒的形状先验。当使用基于实时Transformer的检测器RF-DETR进行评估时,我们的框架在仅使用16GB有限GPU内存的情况下,弥合了跨城市分布差距。我们的优化变体RF-DETR-HR和RF-DETR-Grayworld相较于基线方法取得了+24.29的显著经验增益,在AI City Challenge Track 6排行榜上以47.53 mAP的成绩获得第一名。代码和数据可在以下地址获取:\u003ca href="this https URL"\u003eSKKUAutoLab/aic26_cross_city\u003c/a\u003e。
英文摘要
Real-world deployment of traffic surveillance systems is bottlenecked by geographic domain shift, in which models trained in one city underperform when applied to an unseen target city. Conventional domain adaptation relies on hyperparameter-sensitive architectures or direct profiling of target data. Both are fundamentally precluded in privacy-conscious ecosystems that require completely blind training and evaluation loops. In this setting, we explore the effects of pre-training and augmentation in addressing the domain shift problem. Specifically, we propose a new modular training pipeline for object detection structured around two core orthogonal pillars: (1) a multi-dataset pre-training strategy featuring a class-agnostic objectness distillation to decouple structural vehicle geometry from semantic taxonomies, and (2) a domain-resilient augmentation stream featuring a novel Grayworld transformation that forces global attention heads to strip volatile chromatic shortcuts in favor of robust shape priors. When evaluated with the real-time transformer-based detector RF-DETR, our framework bridges cross-city distribution gaps while using limited GPU memory (16GB). Our optimized variants, RF-DETR-HR and RF-DETR-Grayworld, deliver a substantial empirical gain of +24.29 over the baseline, achieving 1st place (47.53 mAP) on the AI City Challenge Track 6 leaderboard. Code and data are available at: \href{https://github.com/SKKUAutoLab/aic26_cross_city}{SKKUAutoLab/aic26\_cross\_city}.
CommentsThis paper has been accepted by the AI City Challenge Workshop of the European Conference on Computer Vision (ECCV 2026)