arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05583cs.CV

三维视觉-语言模型综述

An overview of 3D Vision-Language Models

Márcus Lobo, Vitor Matias, Afonso Paiva, Jeová Farias, Tiago Novello, Moacir Ponti

首次发表
浏览论文内容

中文总结 AI 辅助

本文综述三维视觉-语言模型,涵盖三维表示编码、跨模态对比对齐及多模态框架,并介绍语言引导的三维高斯泼溅、形状生成和具身AI等最新进展。

中文摘要 AI 辅助

视觉-语言模型(VLMs)正在通过对齐视觉和文本嵌入来重塑计算机视觉领域,使模型能够识别视觉概念并使用自然语言对其进行推理。然而,传统的三维深度学习模型通常针对特定任务(如分类、分割或检测)进行训练,并且天然不支持利用文本或图像作为查询从其嵌入空间中进行跨模态检索。为解决这一问题,基于对比语言-图像预训练(CLIP)的方法将三维嵌入与预训练的图像和文本表示对齐,从而催生了支持三维形状的零样本分类、跨模态检索和开放词汇识别的三维视觉-语言模型(3D VLMs)。本教程概述了三维视觉-语言模型,涵盖从三维表示的基本定义及其嵌入编码,到跨模态对比对齐、现代多模态框架,以及三维视觉-大语言模型(3D VLLMs)等内容。我们介绍了用于多模态嵌入对齐的对比学习的主要定义,并重点介绍了语言引导的三维高斯泼溅、三维形状生成以及面向机器人的具身人工智能方面的最新进展。

英文摘要

Vision-Language Models (VLMs) are reshaping computer vision by aligning visual and textual embeddings, allowing models to recognize visual concepts and reason about them using natural language. Traditional 3D deep-learning models, however, are typically trained for specific tasks, such as classification, segmentation, or detection, and do not naturally support cross-modal retrieval from their embedding spaces using text or images as queries. To address this issue, Contrastive Language-Image Pretraining (CLIP)-based methods align 3D embeddings with pretrained image and text representations, giving rise to 3D Vision-Language Models (3D VLMs) that support zero-shot classification, cross-modal retrieval, and open-vocabulary recognition of 3D shapes. This tutorial provides an overview of 3D VLMs, ranging from basic definitions of 3D representations and their encoding into embeddings to cross-modal contrastive alignment, modern multimodal frameworks, and 3D Vision-Large Language Models (3D VLLMs). We present the main definitions of contrastive learning for multimodal embedding alignment and highlight recent advances in language-guided 3D Gaussian splatting, 3D shape generation, and embodied AI for robotics.

发表机构

  • ICMC-USP(圣保罗大学数学与计算机科学研究所)
  • IMPA(巴西国家纯数学与应用数学研究所)
  • Bowdoin College(鲍登学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑