发表机构
LaSER Lab, Kennesaw State University; Lincoln Institute for Agri-Food Technology, University of Lincoln(激光实验室,肯尼索州立大学; 林肯农业食品技术研究所,林肯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对农业机器人夜间视觉导航中大规模标注图像数据集难获取的问题,提出无监督图像翻译框架,整合CLIP模型并引入可见性掩码,复用白天语义标签训练模型,经实验验证提升了图像质量和下游语义分割性能。
AI 中文摘要
虽然视觉导航在农业机器人领域已得到广泛研究,但大多数现有系统假定为白天条件。夜间部署自主机器人有诸多优势,如24小时作物和土壤监测等。然而,现代基于视觉的系统严重依赖大规模标注图像数据集,夜间操作场景下获取此类数据集颇具挑战。为此,我们提出一种无监督图像翻译框架,可将白天植物行RGB图像转换为近红外(NIR)夜间图像,无需逐像素监督,能直接复用白天语义标签训练夜间感知模型。通过整合预训练的对比语言-图像预训练(CLIP)模型,该框架在日夜翻译中保持语义一致性。此外,引入可见性掩码以考虑夜间场景中NIR照明的有限有效范围。我们与最先进的图像翻译基线进行比较评估,展示了更高的图像质量,在夜间视觉导航的下游语义分割中性能提升。我们利用AgriNight数据集(包含428张白天和549张夜间图像,由配备夜视的移动机器人在农田收集并手动标注逐像素语义标签)进行评估,并将其作为夜间农业视觉导航的首个基准。我们还使用物理机器人进行了夜间实时自主导航实验,并提供了数据和代码链接。
英文摘要
While visual navigation has been extensively studied in agricultural robotics, most existing systems assume daytime conditions. In fact, deploying autonomous robots at night offers significant advantages, including 24-hour crop and soil monitoring, fruit harvesting, and nocturnal pest detection. Modern vision-based systems, however, rely heavily on large-scale well-annotated image datasets, which remains challenging to obtain for nighttime operation scenarios. To address this, we propose an unsupervised image translation framework that converts daytime plant-row RGB images into near-infrared (NIR) nighttime counterparts without requiring pixel-to-pixel supervision. This enables the direct reuse of daytime semantic labels for training nighttime perception models. In particular, by incorporating a pre-trained Contrastive Language-Image Pre-training (CLIP) model, the proposed framework is designed to preserve semantic consistency during day-to-night translation. Additionally, a visibility mask is introduced to account for the limited effective range of NIR illumination in nighttime scenes. We conduct comparative evaluations with state-of-the-art image translation baselines and demonstrate higher image qualities, as supported by improved performance in downstream semantic segmentation for nighttime visual navigation. For evaluation, we utilize AgriNight--a novel dataset comprising 428 daytime and 549 nighttime images collected using night-vision-equipped mobile robots in agricultural fields and manually annotated with pixel-wise semantic labels--and introduce it as the first benchmark for nighttime agricultural visual navigation. We also perform real-time autonomous navigation experiments with a physical robot operating at night. The data and code are available at: https://github.com/mamorobel/AgriNight.
CommentsAccepted to IROS2026