arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过程序合成数据实现番茄表型的文本条件分割

Text-conditioned Segmentation for Tomato Phenotyping via Procedural Synthetic Data

Samy Mounir, Mikolaj Cieslak, Najmeddine Dhieb, Hakim Ghazzai, Jonathan Klein, Katja Froehlich, Soeren Pirk, Wojciech Palubicki, Gianluca Setti, Ahmed M. Eltawil, Dominik L. Michels

arXiv 2607.18576首次发表:更新:

发表机构

Computer, Electrical and Mathematical Sciences and Engineering Division, KAUST; Department of Computer Science, University of Calgary; Department of Computer Science, Kiel University; Department of Artificial Intelligence, Adam Mickiewicz University; Department of Electronics and Telecommunications, Polytechnic University of Turin(计算机、电气与数学科学与工程系,沙特阿卜杜拉国王科技大学; 卡尔加里大学计算机科学系; 基尔大学计算机科学系; 亚当·密茨凯维奇大学人工智能系; 都灵理工大学电子与电信系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对温室作物表型分析中缺乏标注数据问题,提出结合合成数据生成与基础模型微调的框架,通过对商业樱桃番茄温室建模生成合成数据集,微调SAM 3,显著提升分割性能,还公开相关模型与数据集。

AI 中文摘要

基于视觉的自动化是减少温室作物生产和表型分析中人工劳动的理想选择。然而,进展受限于缺乏标注训练数据。基于视觉的基础模型的最新进展在零样本泛化到新领域方面显示出有希望的结果,但在复杂农业环境中性能下降。本文提出了一个用于番茄植株分割的模拟到真实框架,将合成数据生成与基础模型微调相结合。对商业樱桃番茄温室建模,在不同视角、光照条件和植物形态下生成大规模合成数据集。随后在合成数据集上微调Segment Anything Model 3(SAM 3),使其针对温室作物器官的文本条件分割行为专门化,同时保留实现零样本转移的一般视觉先验。通过在多个真实世界温室数据集上评估框架,证明合成数据与SAM 3微调相结合显著提高分割性能和模型置信度。为支持社区基准测试,公开发布程序模型、生成的合成数据集和微调后的SAM 3权重。

英文摘要

Vision-based automation is an excellent candidate for reducing manual labor in greenhouse crop production and phenotyping. However, progress is constrained by the lack of annotated training data. Recent advances in vision-based foundational models have shown promising results in zero-shot generalization to novel domains, but their performance drops in complex agricultural environments. In this work, we present a sim-to-real framework for tomato plant segmentation that combines synthetic data generation with fine-tuning of a foundation model. We model a commercial cherry tomato greenhouse and use it to generate a large-scale synthetic dataset under diverse viewpoints, lighting conditions, and plant morphology. Subsequently, we fine-tune the Segment Anything Model 3 (SAM 3) on the synthetic dataset, specializing its text-conditioned segmentation behavior for greenhouse crop organs while retaining the general visual prior that makes zero-shot transfer possible. By evaluating our framework on multiple real-world greenhouse datasets, we demonstrate that combining synthetic data with SAM 3 fine-tuning significantly improves segmentation performance and model confidence. To support community benchmarking, we publicly release the procedural model, the generated synthetic dataset, and our fine-tuned SAM 3 weights.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑