arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UAV3DCrop:面向重复多视角无人机作物调查的三维重建基准测试

UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys

Junxiong Zhou, Xuechen Li, Chonghao Qiu, Lang Qiao, Xiaowei Jia, Qi Yang, Chishan Zhang, Leikun Yin, Nanshan You, Vipin Kumar, David Mulla, Ce Yang, Zhenong Jin, Licheng Liu

arXiv 2608.06404首次发表:更新:

发表机构

University of Minnesota, Twin Cities; University of Wisconsin–Madison; University of Pittsburgh; Max Planck Institute for Biogeochemistry; Boston University; Peking University(明尼苏达大学双城校区; 威斯康星大学麦迪逊分校; 匹兹堡大学; 马克斯·普朗克生物地球化学研究所; 波士顿大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出UAV3DCrop基准,评估NeRF、3DGS等场景优化方法及预训练前馈模型在作物三维重建任务的表现,发现无单一方法适配所有农艺需求,仅一种前馈模型可恢复可用度量尺度。

AI 中文摘要

准确的作物三维监测是数据驱动精准农业的基础,可实现田间尺度下对植物结构、生长动态及管理响应的分析。现代三维重建方法在通用基准上表现出色,但生成的外观可能无法转化为作物田区具有度量学和农艺学价值的几何结构。本文提出UAV3DCrop,这是一个公开的重复多视角无人机(UAV)作物调查基准,包含来自91个场景的88830张RGB图像,像素分辨率为5280×3956,地面采样距离为3.6-5.8 mm,场景覆盖玉米、大豆、小麦和燕麦。Track A针对保留视图、摄影测量参考深度及冠层高度恢复,评估了7种场景优化方法——神经辐射场(NeRF)和三维高斯溅射(3DGS)变体;Track B测试了4种预训练前馈模型的零样本相机位姿与几何估计能力。场景优化方法在三个目标任务上排名不同:Splatfacto-big在外观任务中领先,而Scaffold-GS在深度任务中领先,且与Splatfacto在冠层高度任务上表现无统计学差异。在前馈模型中,MapAnything在8项指标中的7项上领先,其余模型在不同作物间表现差异更大,且在绝对尺度上存在严重缺陷,仅通过对齐操作掩盖了该问题。重复采集揭示了与输出类型、模型相关的进一步敏感性,这些敏感性与采集序列中的位置及 tie-point 多重性相关。因此,当前三维重建方法暂不适用于农艺用途:没有单一方法能同时在外观、几何和冠层高度任务中胜出,且4种前馈模型中仅有一种能恢复可用的度量尺度。该数据集可通过此https URL公开获取。

英文摘要

Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and management response. Modern 3D reconstruction methods perform strongly on generic benchmarks, but rendered appearance may not translate into metrically and agronomically useful geometry in crop fields. We introduce UAV3DCrop, a public benchmark of repeated multi-angle unmanned aerial vehicle (UAV) crop surveys. It contains 88,830 RGB images at $5280 \times 3956$ pixels, with a ground sampling distance of 3.6-5.8 mm, from 91 scenes spanning corn, soybean, wheat, and oat. Track A evaluates seven scene-optimized methods -- Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) variants -- on held-out views, photogrammetry-referenced depth, and canopy-height recovery. Track B tests four pretrained feed-forward models on zero-shot camera-pose and geometry estimation. The scene-optimized methods rank differently across the three targets: Splatfacto-big leads appearance, whereas Scaffold-GS leads depth and is statistically tied with Splatfacto for canopy height. Among feed-forward models, MapAnything leads on seven of the eight metrics, while the remaining models vary more across crops and fail severely on absolute scale in a way that alignment conceals. Repeated acquisitions reveal further sensitivities that differ by output type and by model, associated with position within the acquisition sequence and with tie-point multiplicity. Current 3D reconstruction methods are therefore not yet interchangeable for agronomic use: no single method wins on appearance, geometry, and canopy height at once, and only one of four feed-forward models recovers usable metric scale. The dataset is publicly available at https://link-dev.github.io/UAV3DCrop/

Comments22 pages, 7 figures. Dataset and project page: https://link-dev.github.io/UAV3DCrop/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑