arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17559cs.ROcs.AIcs.ETcs.LG

COLIP-2:嗅觉-视觉-语言嵌入

COLIP-2: Olfaction-Vision-Language Embeddings

Kordel Kade France

首次发表
浏览论文内容

中文总结 AI 辅助

COLIP-2模型构建多模态嵌入空间,将嗅觉与视觉、语言结合,通过训练使机器人能定位香气来源。因缺乏相关大规模数据集,需新方法和数据集。文中展示其内部测试结果并优化,该模型有望用于多模态领域。

中文摘要 AI 辅助

对比嗅觉-语言-图像预训练2(COLIP-2)模型是一个多模态嵌入空间,将嗅觉置于视觉和语言中作为一等公民。分子结构、气体传感器读数、气味描述语言和图像都被训练到一个共享表示空间,使机器人能概率性地将检测到的香气定位到场景中的物体。由于没有配对图像-气味示例的ImageNet规模数据集,所以需要收集。发布COLIP-2的目的是展示利用开源嗅觉数据为机器人构建的能力极限,论证为何需要新方法和数据集来实现先进的嗅觉感知能力。文中列举了COLIP-2架构内部测试结果并进行必要优化以在边缘运行模型用于实时机器人应用。COLIP-2虽为机器人设计,但受多学科专家影响,希望该模型能在任何需要嗅觉智能的多模态领域有用。

英文摘要

The Contrastive Olfaction-Language-Image Pre-training 2 (COLIP-2) model is a multimodal embeddings space that places olfaction as a first-class citizen among vision and language. Molecular structure, gas-sensor readings, odor-descriptor language, and images are all trained into a single shared representation space, so that a robot can localize a detected aroma to objects in a scene probabilistically. No ImageNet-scale datasets of paired image-scent examples exists which warrants the need for their collection. Our intent with the release of COLIP-2 is to demonstrate the limit of what can be built for robotics with open-sourced olfactory data in order to ground the argument for why new methodologies and datasets are necessary in order to enable advanced olfactory-oriented perception capabilities. We enumerate results from internal testing of the COLIP-2 architecture and make necessary optimizations to run the model at the edge for real-time robotics applications. While developed with robotics in mind, the design of COLIP-2 has been influenced by experts across many disciplines of science in academia and industry, and we hope that the model can be useful in any multimodal domain requiring olfactory intelligence.

发表机构

  • Scentience

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑