建筑环境视觉-语言表征中效价与唤醒度的组织:来自EMOIS数据集的洞见
Organization of Valence and Arousal in Vision-Language Representations of Built Environments: Insights from the EMOIS Dataset
浏览论文内容
中文总结 AI 辅助
本文提出EMOIS数据集,含1,544张建筑环境图像及效价和唤醒度评分,利用CLIP表征分析发现效价组织更强,回归预测性能高,为情感计算提供资源。
中文摘要 AI 辅助
建筑环境的视觉感知有助于形成人们在日常生活中产生的情感印象。然而,这些印象如何在视觉基础模型中被表征仍 largely 未被探索。为支持对该主题的系统性研究,我们引入了情感空间印象(EMOIS)数据集,包含1,544张真实世界的建筑环境图像。每张图像都标注了由日本成年人通过大规模网络调查收集的图像诱发效价和唤醒度评分,每张图像约有120个评分。利用对比语言-图像预训练(CLIP)表征,我们进行了预测性和几何分析,以系统性地研究效价和唤醒度如何在表征空间中被编码和组织。这些分析揭示,效价表现出比唤醒度更强且更连贯的组织。与开放情感标准化图像集(OASIS)——一个通用情感照片的基准数据集——进行的跨数据集分析揭示了两数据集之间情感组织的差异。回归分析展示了在EMOIS内对效价和唤醒度的高预测性能,在重复的内部留出评估中,平均决定系数分别为0.865和0.807。最后,我们展示了一个基于示例的界面,说明学习到的表征如何支持对预测情感值的定性解释。这些发现有助于阐明建筑环境的情感表征,并将EMOIS确立为该领域未来情感计算研究的密集标注资源。
英文摘要
Visual perception of built environments contributes to the affective impressions that people form in everyday life. However, how these impressions are represented within vision foundation models remains largely unexplored. To support the systematic investigation of this subject, we introduce the Emotional Impression of Spaces (EMOIS) dataset, comprising 1,544 real-world built-environment images. Each image is annotated with image-evoked valence and arousal ratings collected from Japanese adults by conducting a large-scale web-based survey, with approximately 120 ratings per image. Using Contrastive Language--Image Pre-training (CLIP) representations, we perform predictive and geometric analyses to systematically investigate how valence and arousal are encoded and organized within the representation space. These analyses reveal that valence exhibited stronger and more coherent organization than arousal. Cross-dataset analyses with the Open Affective Standardized Image Set (OASIS), a benchmark dataset of general affective photographs, reveal differences in affective organization between the two datasets. Regression analyses demonstrate high predictive performance for valence and arousal within EMOIS, with mean coefficients of determination of 0.865 and 0.807, respectively, across repeated internal hold-out evaluations. Finally, we present an example-based interface illustrating how learned representations can support qualitative interpretation of predicted affective values. These findings can help elucidate affective representations of built environments and establish EMOIS as a densely annotated resource for future affective computing research in this domain.
发表机构
- Toyota Central R&D Labs., Inc.(丰田中央研发实验室有限公司)
机构由 AI 辅助整理,请以论文原文为准。