发表机构
Swiss Federal Institute of Technology Lausanne (EPFL)(瑞士洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用跨模态学习机制,提出在特定测试环境下的受限设置并开发TST方法,通过在测试环境收集多模态数据自监督预训练模型,评估其在下游任务的表现,发现可减少对外部大规模数据集预训练的依赖。
AI 中文摘要
跨模态学习是通过利用多模态进行自监督的基本机制,许多实际应用涉及在测试环境中进行多模态传感的设备。发育心理学研究表明生物也利用它构建周围环境的有效表征。为研究此,作者提出一种受限设置,开发了测试空间训练(TST),在测试环境中收集多模态数据并进行自监督预训练,在相同环境的下游任务上评估模型。结果发现仅从测试环境收集丰富多模态数据并利用跨模态学习,能与在大规模互联网数据集上预训练的通用模型取得有竞争力的结果,还进行了相关分析和消融实验。
英文摘要
Cross-modal learning, i.e., learning to predict one modality from another, is a fundamental mechanism for self-supervision via leveraging multimodality. Many practical applications, e.g., deploying a household robot, involve devices that are equipped with a rich set of sensors that enable multimodal sensing in their test environment. This presents an opportunity to apply cross-modal learning to the multimodal data sensed by these devices to learn representations. Findings in developmental psychology also suggest that biological agents leverage it to build an effective representation of their surroundings. To study this, we propose a controlled setup, where we restrict a user device to just a given test environment. It results in a specialization setup where we attempt to develop a performant model for this specific test environment. Under this setup, we develop Test-Space Training (TST), which performs multimodal data collection in the test environment and performs self-supervised pre-training on it. We evaluate these models on various downstream tasks in the same environment. Under this setup, we find various interesting insights, such as collecting rich multimodal data only from the test environment and leveraging cross-modal learning, we can achieve competitive results with generalist models (e.g., DINOv2 and CLIP) pre-trained on large-scale internet datasets. This enables an alternative scenario where the need for external Internet-scale datasets for pre-training models is reduced. We also present a set of analyses and ablations that raise intriguing points on substituting data with (multi)modality, and how varying pre-training data enables a tradeoff between a model's abilities to specialise to a test environment, and generalize to held-out spaces.
CommentsProject page: https://tst-vision.epfl.ch/