发表机构
University of California Santa Cruz(加州大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出BirdsEye,一种人在回路地理空间标注方法,利用RTK定位和射影几何将专家标注从图像转移到现场,实现快速数据集构建,在农业场景中显著提升标注效率并保持亚分米级精度。
AI 中文摘要
真实世界的感知系统必须适应不断变化的环境,但人工图像标注无法扩展到现场数据规模。我们提出了BirdsEye,它将专家标注从图像转移到现场:操作员使用RTK定位记录世界坐标中的目标位置,并通过校准的射影几何将每次观测传播到目标可见的所有帧。为了量化物理标注与图像观测的对齐程度,我们推导了从相机位姿不确定性到像素不确定性的第一阶映射,并通过蒙特卡洛模拟进行了验证。该映射在六个每轴位姿方差上是线性的,因此可逆转为传感器设计工具:我们给出了将标注容差转换为可容许位姿噪声预算的凸集的充分条件,一个部署传感器套件的最大可容许缩放的闭式解,以及在等预算份额分配下的唯一每轴位姿规格。我们还分析了投影所依据的平面表面近似,该近似在地形坡度不超过10度时成立。通过直接测量,我们表明在排除持续偏航的条件下,系统投影精度在AGL高度10-20米处为亚分米级(亚30像素)。在三个农业场地的现场案例研究中,两名现场工作人员在大约12小时内产生了12,524个标注帧,携带55,600个标签(每位工人标注速率比人工标注提高25.5倍)。使用此工作流程收集的图像训练的检测器,在预先注册的操作点上,在地理位置不同的农场中恢复了视野内调查目标的56-89%;对领先配置的人工审查估计检测精度为83-87%,涵盖三种针对携带矛盾人工裁决的聚类的决胜约定。
英文摘要
Real-world perception systems must adapt to changing environments, but manual image annotation cannot scale to field data volumes. We present BirdsEye, which shifts expert annotation from images to the field: an operator records target locations in world coordinates using RTK positioning and calibrated projective geometry propagates each observation to all frames where the target is visible. To quantify how well physical annotations align with image observations, we derive a first-order mapping from camera-pose uncertainty to pixel uncertainty and validate it against Monte Carlo simulation. This mapping is linear in the six per-axis pose variances, so it inverts into a sensor design tool: we give a sufficient condition converting an annotation tolerance into a convex set of admissible pose-noise budgets, a closed-form largest admissible scaling of a deployed sensor suite, and a unique per-axis pose specification under an equal-budget-share allocation. We also analyze the planar-surface approximation underlying the projection, which holds up to 10 degrees of terrain slope. By direct measurement, we show that system projection accuracy is sub-decimeter (sub-30 pixel) at AGL altitudes of 10-20m under conditions excluding sustained yawing. During an in-field case study across three agricultural sites, two field workers produced 12,524 annotated frames carrying 55,600 labels in roughly 12 hours (25.5x per-worker rate increase over manual labeling). Detectors trained on imagery collected by this workflow recovered 56-89% of in-view surveyed targets at a geographically distinct farm, at pre-registered operating points; human review of the leading configuration estimates detection precision at 83-87%, spanning three tie-break conventions for clusters carrying contradictory human verdicts.
Comments30 pages, 10 figures (plus 2 in appendix)