arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于跨城市目标检测的冻结高分辨率推理:2026年AI城市挑战赛研究

Frozen High-Resolution Inference for Cross-City Object Detection: An AI City Challenge 2026 Study

Jaeuk Kim

arXiv 2608.03136首次发表:更新:

AI 中文总结

本研究在2026年AI城市挑战赛中,针对跨城市目标检测任务,发现对704×704训练的RF-DETR-Large检测器进行冻结1120×1120推理可提升聚合AP,同时警示域内验证在仅聚合反馈下不可靠。

AI 中文摘要

跨城市目标检测要求在一个城市训练的检测器能够泛化到未标记的目标城市。在2026年AI城市挑战赛第6赛道中,我们分析了孤立训练即服务(Training-as-a-Service)平台内单个RF-DETR-Large检测器的存档配置,该平台的服务器仅返回源城市与目标城市图像的隐藏混合数据集上的聚合COCO风格平均精度(AP)。在704×704分辨率下训练的检查点进行冻结的1120×1120推理,在所有评估配置中实现了最高的聚合AP(从0.3272提升至0.3654,增幅0.0382),且无需任何参数更新;该设置在输入像素量为2.53倍时,对小物体的相对增益最大,对中物体的绝对增益最大。一种热启动的1120像素微调方案达到了0.3470的AP,但其域内验证AP从0.767升至0.789,这警示在仅聚合的跨城市反馈下,域内验证并非可靠的模型选择信号。由于该运行的评估使用了比仅推理运行更高的置信度阈值(0.05对比0.01),我们将其得分视为描述性存档结果,而非对微调的受控判定。灰度世界归一化未显著改变冻结1120的结果,且经核查发现某一矩形运行使用了意外的纵向方向。我们发布了逐字平台命令、配置快照及每项结论的明确证据边界。每项配置仅提交一次,最佳配置由隐藏服务器选定,因此这些是关于聚合混合的探索性、可核查发现,未确立目标城市特定的改进。

英文摘要

Cross-city object detection requires a detector trained in one city to generalize to an unlabeled target city. In AI City Challenge 2026 Track 6, we analyze archived configurations of a single RF-DETR-Large detector inside an air-gapped Training-as-a-Service platform whose server returns only an aggregate COCO-style AP over a hidden mixture of source- and target-city images. Frozen 1120 x 1120 inference of a checkpoint trained at 704 x 704 achieved the highest aggregate AP among the evaluated configurations (0.3272 -> 0.3654, +0.0382) without any parameter update, with the largest relative gain on small objects and the largest absolute gain on medium objects, at 2.53x the input pixels. A warm-start 1120px fine-tuning recipe reached 0.3470 while its in-domain validation AP rose (0.767 -> 0.789), a caution that in-domain validation is an unreliable model-selection signal under aggregate-only cross-city feedback. Because that run's evaluation used a higher confidence threshold than the inference-only runs (0.05 vs. 0.01), we treat its score as a descriptive archived outcome rather than a controlled verdict on fine-tuning. Gray-world normalization did not meaningfully change the frozen-1120 result, and a rectangular run was found by audit to have used an unintended portrait orientation. We release verbatim platform commands, configuration snapshots, and an explicit evidence boundary for every claim. Each configuration was submitted once and the best was selected on the hidden server, so these are exploratory, audited findings about the aggregate mixture; they do not establish target-city-specific improvement.

Comments16 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑