表象可能具有欺骗性:众包航空损伤评估中不同影像来源下标注者与审核者的表现
Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment
- Texas A&M University(德克萨斯农工大学)
- University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究多源航空影像损伤评估的标注与审核表现,发现低分辨率影像标注分歧更高,为多源数据集构建提出三项建议。
AI中文摘要:
本文开展了首个已知的、针对多源遥感影像下标注者与审核者表现的实证研究,评估了无人机、有人航空及卫星影像的人工标注情况。由于现有航空影像数据集主要依赖单源影像,目前尚无公认的实践标准来高效分配人力以构建大规模多源航空数据集。本研究通过分析9场灾害的灾后建筑物损伤评估数据集中标注者与审核者的表现,解决了这一局限:该数据集包含无人机影像中的20041栋建筑物、有人航空影像中的20695栋建筑物、卫星影像中的33392栋建筑物,由187名标注者完成标注,随后经过两轮质量控制阶段:单轮审核者审核及共识委员会审核。分析得出两项对标准众包实践提出疑问的发现:其一,最终委员会对初始标注的修正率从高分辨率到低分辨率来源急剧上升(有人航空为25.27%,卫星为36.95%),且在所有观测到的工作流程阶段均保持该排序;其二,单轮个人审核虽减少了分歧但未解决分歧:审核后,委员会仍修正了无人机影像标注的6.85%、有人航空影像标注的14.05%、卫星影像标注的20.86%。这些观察表明,在类似本研究的工作流程中,统一的审核分配会在低分辨率影像中留下最多的剩余分歧。基于此证据,且与自适应任务分配、预算感知质量控制的现有研究一致,本文为多源数据集构建提供三项建议。
英文摘要:
This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed imagery, evaluating human labeling across drone, crewed aviation, and satellite views. Because existing aerial imagery datasets rely predominantly on single-source imagery, there is no currently established state of practice for efficiently allocating human labor to curate large-scale, multi-source aerial datasets. This work addresses this limitation by analyzing annotator and reviewer performance within a post-disaster building damage assessment dataset of 9 disasters, where 20041 buildings in drone, 20695 buildings in crewed aviation, and 33392 buildings in satellite imagery were labeled. These labels, provided by 187 annotators, were then refined through two successive quality-control stages: a single-reviewer pass followed by a consensus-committee review. Our analysis reveals two findings that raise questions for standard crowd-sourcing practices. First, initial annotations were revised by the final committee at rates that rise steeply from higher- to lower-resolution sources (25.27% for crewed aviation and 36.95% for satellite), with the same ordering at every observed workflow stage. Second, a single individual review reduced but did not resolve this disagreement: after review, the committee still revised 6.85% of drone, 14.05% of crewed, and 20.86% of satellite labels. These observations suggest that, in workflows like this one, uniform review allocation leaves the most residual disagreement in lower-resolution imagery. Based on this evidence, and consistent with prior work on adaptive task assignment and budget-aware quality control, this paper offers three recommendations for multi-source dataset curation.