跨多样海岸环境的视觉-语言模型评估
Evaluation of Vision-Language Models Across Diverse Coastal Environments
浏览论文内容
中文总结 AI 辅助
本研究提出密集标注的海岸数据集,评估七个视觉-语言模型,发现海岸类别性能较低主要受分割和语言表示影响,而非环境本身。
中文摘要 AI 辅助
视觉-语言模型(VLMs)通过将视觉观察与自然语言概念关联来实现机器人感知。然而,它们在海岸环境中的性能在很大程度上尚未被探索。我们引入了一个密集标注的海岸数据集,包含在夏威夷欧胡岛三个地区的七次任务中收集的超过1,000张图像,涵盖18个语义类别和超过7,400个标注实例。我们通过三个互补实验评估了七个现代视觉-语言模型,分别测量文本到掩码、掩码到掩码以及掩码到文本的对齐。总体而言,广阔景观类别比传统物体和海岸类别识别得更准确,其中海岸概念构成了最大的挑战。然而,对海岸和陆地数据集中共享的传统类别的比较显示,没有一致性的性能差异可仅归因于环境背景。掩码到掩码匹配在传统和海岸类别中也保持相似,而替代性文本标签显著提高了对若干海岸概念的识别。这些结果表明,海岸类别上较低的性能(至少在所评估的物体/查询类别上)在很大程度上受到分割和语言表示的影响。
英文摘要
Vision-language models (VLMs) enable robotic per- ception by associating visual observations with natural-language concepts. Yet their performance in coastal environments remains largely unexplored. We introduce a densely labeled coastal dataset containing more than 1,000 images collected across seven missions in three regions of Oahu, Hawaii, with 18 semantic classes and over 7,400 annotated instances. We evaluate seven modern VLMs through three complementary experiments mea- suring text-to-mask, mask-to-mask, and mask-to-text alignment. Broad landscape classes are generally recognized more accurately than conventional object and coastal classes, with coastal con- cepts presenting the greatest challenge. However, comparisons of shared conventional classes across coastal and terrestrial datasets reveal no consistent performance difference attributable solely to environmental context. Mask-to-mask matching also remains similar across conventional and coastal classes, while alternative textual labels substantially improve recognition of several coastal concepts. These results suggest that lower performance on coastal classes (at least on the objects/query categories evaluated) is heavily influenced by segmentation and linguistic representation.
发表机构
- Brigham Young University(杨百翰大学)
机构由 AI 辅助整理,请以论文原文为准。