AI 中文总结
针对农业场景可通行性模糊问题,提出TEA模块并集成至AgriVLN构建TEA-AgriVLN,在A2A基准上将SR提至0.54、NE降至2.70米,实现农业VLN-CE领域最优性能。
AI 中文摘要
连续环境下的视觉语言导航(VLN-CE)要求智能体遵循自然语言指令,预测一系列低级动作以操控机器人从起点导航至目标位置。A2A基准与AgriVLN方法率先将VLN-CE从室内场景扩展至农业场景,然而研究发现二者存在关键差异:室内场景中区域可通行性分类通常明确,如木地板可通行、混凝土墙不可通行;但农业场景中该问题往往模糊,例如未成熟玉米田对机器狗可能可通行,对人类则可能不可通行。为解决该问题,本文提出TEA模块,其功能为估计相机图像的可通行性,当预测动作与可通行性地图不匹配时,向决策模块发送告警以触发重新思考。将该模块集成至AgriVLN主干网络,构建TEA-AgriVLN方法。在A2A基准上评估时,该方法将成功率(SR)从0.47提升至0.54,导航误差(NE)从2.91米降至2.70米,在农业VLN-CE领域实现了当前最优性能。本文还开展了消融研究与案例研究,探讨了TEA在不同地面类别和场景类别上的有效性与局限性。代码:this https URL。
英文摘要
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow a natural language instruction, predicting a sequence of low-level actions to navigate a robot from a starting point to a target location. The A2A benchmark and the AgriVLN method pioneeringly extended VLN-CE from indoor scenes to agricultural scenes, while we observed a challenging distinction: In indoor scenes, whether a zone is traversable tends to be clear to classify, such as wood floors are traversable but concrete walls are not. In agricultural scenes, however, this issue tends to be ambiguous, such as an unripe cornfield might be traversable for a robotic dog but might be non-traversable for a human. To address this issue, we propose the TEA module, which estimates the traversability of the camera image, then alarm the decision-maker for rethinking when the predicted action does not align with the traversability map. We integrate it into the AgriVLN backbone to build our TEA-AgriVLN method. When evaluated on A2A, it improves Success Rate (SR) from 0.47 to 0.54 and Navigation Error (NE) from 2.91 m to 2.70 m, showing the state-of-the-art performance in the agricultural VLN-CE domain. We further implement the ablation studies and the case study, discussing the effectiveness and limitations of TEA on different ground categories and scene classes. Code: https://github.com/AlexTraveling/TEA-AgriVLN.