基于视觉语言模型(VLM)的田间玉米点云端到端程序化建模流程
A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds
- Iowa State University(爱荷华州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出一种基于VLM的自动化流程,无需人工调整或物种特定训练数据,即可从点云重建程序化玉米3D模型,在100株田间玉米上达到5.4毫米中位Chamfer距离,并恢复99.4%的参考叶片。
AI中文摘要:
可编辑的田间作物3D模型支持高通量表型分析和计算机模拟育种试验,但从扫描点云构建这些模型需要器官级分割和拟合。程序化生成器可以将器官级参数集转换为适合分析的3D模型,但获取该参数集需要每株植物数小时的人工调整或基于物种特定标签训练的分割模型。我们提出了一种自动化流程,无需人工调整或物种特定训练数据,即可从原始3D点云重建程序化玉米模型。多模态视觉语言模型(VLM)在渲染的正交视图中标注叶片中线。确定性几何算法将标注反投影到点云上,通过跨视图共识将其合并为3D叶片,并在基于方向加权的表面图上将中线扩展为完整叶片。测量的器官参数填充基于非均匀有理B样条(NURBS)的程序化模型生成器的植物描述符。然后通过可微NURBS拟合,根据扫描点细化每个叶片表面。该流程在来自MaizeField3D数据集的100株基因型多样的田间玉米植株上,达到了5.4毫米的中位整株Chamfer距离。重建结果比早期基于人工标注的半自动流程更接近扫描数据。该流程在不使用这些标签作为输入的情况下,恢复了1,023个精选参考叶片中的1,017个(99.4%),交并比至少为0.5。这些结果表明,当下游几何阶段能够纠正VLM标注时,VLM标注可成为可用的器官级测量数据。这使得在现代表型实验规模下自动生成可编辑的3D植物资产成为可能。
英文摘要:
Editable 3D models of field-grown crops support high-throughput phenotyping and in silico breeding trials, but building them from scanned point clouds requires organ-level segmentation and fitting. Procedural generators can turn an organ-level parameter set into an analysis-suitable 3D model, but obtaining that set requires hours of manual tuning per plant or segmentation models trained on species-specific labels. We present an automated pipeline that reconstructs procedural maize models from raw 3D point clouds without manual tuning or species-specific training data. A multimodal vision-language model (VLM) annotates leaf midlines in rendered orthographic views. Deterministic geometric algorithms back-project the annotations onto the point cloud, merge them into 3D leaves by cross-view consensus, and grow the midlines to full blades on an orientation-weighted surface graph. Measured organ parameters populate a plant descriptor for a Non-Uniform Rational B-Spline (NURBS)-based procedural model generator. Each leaf surface is then refined against its scan points by differentiable NURBS fitting. The pipeline reached a median whole-plant Chamfer distance of 5.4 mm on 100 genotypically diverse field-grown maize plants from the MaizeField3D dataset. The reconstructions were closer to the scans than those of an earlier semi-automated pipeline based on manual annotations. The pipeline recovered 1,017 of 1,023 (99.4%) curated reference leaves at an intersection-over-union of at least 0.5 without using those labels as input. These results show that VLM annotations become usable organ-level measurements when downstream geometric stages can correct them. This makes automated generation of editable 3D plant assets feasible at the scale of modern phenotyping experiments.