arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

全景街道分割的视觉Transformer集成

Vision Transformer Ensembles for Panoramic Street Segmentation

Yunus Serhat Bıçakçı

arXiv 2610.06063首次发表:更新:

发表机构

Marmara University; University of Glasgow(马尔马拉大学; 格拉斯哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于视觉Transformer集成的全景街道分割方法,通过等权平均多个仅编码器掩码Transformer模型的预测,在PalmCity挑战赛中取得第一名,实现60.95%的平均交并比。

AI 中文摘要

街道全景图的语义分割能够支持对城市环境的详细描述,然而小数据集和不平等的训练成本使得模型选择变得困难。本文介绍了在2026年10月5日排行榜快照中PalmCity挑战赛第一名提交所使用的系统。九个预训练分割系统在近似相等的计算预算下进行了比较。候选模型包括DeepLabV3+、SegFormer、UPerNet、Mask2Former、带有线性解码器的DINOv3,以及使用DINOv3的仅编码器掩码Transformer。两个领先的候选模型使用三个随机种子和更长的预算独立训练。对三个仅编码器掩码Transformer模型的类别概率进行等权平均,并在三个图像尺度上评估且结合水平反射,在84张图像的公开验证集上产生了60.95%的平均交并比和71.16%的平均F1分数。提交的预测在隐藏测试排行榜上获得了57.08%的平均交并比和67.96%的平均F1分数。生成全部249个测试掩码耗时251.49秒,包括在一张NVIDIA RTX 5090上的模型初始化和来源检查。峰值分配的GPU内存为2.70 GiB。该研究报告了所有符合条件的模型、所有推理变体、类别级错误、来源条件和可复现性检查,提供了使用现有架构的有文档记录的挑战工作流程。

英文摘要

Semantic segmentation of street panoramas can support detailed descriptions of urban environments, yet small datasets and unequal training costs make model selection difficult. This paper presents the system used for a first place submission to the PalmCity challenge in the leaderboard snapshot dated 5 October 2026. Nine pretrained segmentation systems are compared using approximately equal computation budgets. The candidates include DeepLabV3+, SegFormer, UPerNet, Mask2Former, DINOv3 with a linear decoder, and an Encoder only Mask Transformer using DINOv3. The two leading candidates are trained independently with three random seeds and longer budgets. Equal averaging of class probabilities from the three Encoder only Mask Transformer models, evaluated at three image scales with horizontal reflection, produces 60.95% mean intersection over union and 71.16% mean F1 on the 84 image public validation split. The submitted predictions receive 57.08% mean intersection over union and 67.96% mean F1 on the hidden test leaderboard. Producing all 249 test masks takes 251.49 seconds including model initialization and provenance checks on one NVIDIA RTX 5090. Peak allocated GPU memory is 2.70 GiB. The study reports all eligible models, all inference variants, class level errors, source conditions, and reproducibility checks, providing a documented challenge workflow with existing architectures.

Comments18 pages, 4 figures. Code available at https://github.com/yunusserhat/palmcity_challenge . Trained models available at https://huggingface.co/yunusserhat/palmcity-eomt-dinov3-large

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑