arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoSA:基于相关性的变化注意力与可学习残差门控用于遥感变化检测

Data-Efficient Crosswalk Segmentation from Overhead CCTV via Confidence- and Geometry-Guided Pseudo-Labeling

Abdirashid Omar, Jonghyuk Park

arXiv 2609.08914首次发表:更新:

发表机构

Graduate School of Kookmin University(崇实大学研究生院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对固定交通摄像头斑马线检测的视角偏移问题,提出数据高效目标域流程,结合置信度与几何过滤,在40张人工验证图像上达88.91% IoU,并澄清伪标签评估需与自训练标签隔离。

AI 中文摘要

固定交通摄像头图像的像素级标注成本高昂,而从街景图像训练的斑马线模型在应用于高位CCTV时面临显著的视角和外观偏移。我们研究了一种数据高效的目标域流程,使用了241张人工标注的CCTV图像和5,926个未标注的CCTV帧。源域实验在3,300张第一人称视角(FPV)图像上训练了一个31.0M参数的定制U-Net,并在其330张图像的FPV测试集上获得了93.05%的IoU。该结果是源域基线,而非迁移性能:发布的CCTV笔记本实例化了从torchvision权重初始化的42.0M参数DeepLabV3-ResNet50,且未实现与U-Net检查点的兼容映射。在201张人工标注CCTV图像上训练并在40张留出的人工掩膜上选择,获得了88.91%的IoU。随后模型预测所有未标注帧;图像级置信度和最大连通分量面积先验对候选进行排序,前1,000个候选的平均置信度为0.976,平均综合得分为0.988。仓库审计显示,报告的第二阶段98.52%的IoU是在包含仅由教师生成的伪掩膜的150张图像分割上测量的。由于目录布局不匹配,执行的组合数据加载器未找到任何人工样本,并将1,000个伪标签样本分为850个训练样本和150个评估样本。因此,我们将98.52%报告为内部伪标签一致性,而非人类真实标注准确率。可辩护的目标域结果是40张人工验证图像上的88.91% IoU。在NVIDIA RTX A6000 48 GB GPU上,批次大小为1的FP32推理在512 x 512分辨率下需要12.98毫秒,对应77.03 FPS。这些发现支持了置信度和几何过滤的实用性,同时也说明了为什么伪标签评估必须与用于自训练的标签保持隔离。

英文摘要

Pixel-level annotation of fixed traffic-camera imagery is expensive, while crosswalk models trained from street-level imagery face a substantial viewpoint and appearance shift when applied to elevated CCTV. We investigate a data-efficient target-domain pipeline using 241 manually annotated CCTV images and 5,926 unlabeled CCTV frames. A source-domain experiment trains a 31.0M-parameter custom U-Net on 3,300 first-person-view (FPV) images and obtains 93.05% IoU on its 330-image FPV test split. This result is a source baseline, not transferred performance: the released CCTV notebook instantiates a 42.0M-parameter DeepLabV3-ResNet50 from torchvision weights, and no compatible mapping from the U-Net checkpoint is implemented. Training on 201 manual CCTV images and selecting on 40 held-out manual masks yields 88.91% IoU. The model then predicts all unlabeled frames; image-level certainty and a largest-component area prior rank the candidates, and the top 1,000 attain mean certainty 0.976 and mean combined score 0.988. A repository audit shows that the reported second-stage 98.52% IoU was measured on a 150-image split containing only teacher-generated pseudo-masks. Because of a directory-layout mismatch, the executed combined-data loader found zero manual samples and split 1,000 pseudo-labeled samples into 850 training and 150 evaluation samples. We therefore report 98.52% as internal pseudo-label agreement rather than human-ground-truth accuracy. The defensible target-domain result is 88.91% IoU on the 40 manual validation images. Batch-one FP32 inference at 512 x 512 requires 12.98 ms, corresponding to 77.03 FPS, on an NVIDIA RTX A6000 48 GB GPU. These findings support the practicality of confidence-and-geometry filtering while also showing why pseudo-label evaluation must remain isolated from the labels used for self-training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑