发表机构
NVIDIA; University of Ottawa(英伟达; 渥太华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出多模态数据语言表示框架,将观察结果表示为原子命题,通过全局语义码本统一为共享词汇表,置于可解释空间,实现跨模态理解等,还在自动驾驶等数据上进行了展示。
AI 中文摘要
我们提出了一种用于多模态数据的语言表示,其中任何观察结果,无论是图像、视频还是文本,都被表示为一袋原子命题,即关于场景中实体、动作和关系的简单陈述。全局语义码本将这些统一为规范原子命题的共享词汇表,将每个模态和观察结果置于一个可解释的空间中,该空间跨越细粒度事实到高级概念,并组合成更丰富的概念。这带来了推理的可解释性、跨模态理解和检索以及组合性,从而实现复杂的多模态理解、丰富的数据管理和复杂的结构化检索。我们在自动驾驶和开放世界数据上展示了该框架。
英文摘要
We propose a language representation for multimodal data in which any observation, whether image, video, or text, is expressed as a bag of atomic propositions, simple statements about the entities, actions, and relations in a scene. A global semantic codebook unifies these into a shared vocabulary of canonical atomic propositions, placing every modality and observation into one interpretable space that spans fine grained facts to high level concepts and composes into richer ones. This brings interpretability with reasoning, cross-modal understanding and retrieval, and compositionality that enables complex multimodal understanding, rich data curation and complex structured retrieval. We demonstrate the framework on autonomous driving and open-world data.