情感概念在视觉-语言模型中能否跨来源、模态和架构泛化?
Do Emotion Concepts Generalize Across Sources, Modalities, and Architectures in Vision-Language Models?
浏览论文内容
中文总结 AI 辅助
本研究构建CMES数据集,探究情感概念在视觉-语言模型中跨来源、模态和架构的泛化,发现情感表征共享关系结构和因果效应,即使向量方向不同。
中文摘要 AI 辅助
近期研究表明,大型语言模型将情感概念编码为结构化的内部表征,但现有工作大多聚焦于文本和单一架构。因此,我们提出疑问:情感概念在视觉-语言模型(VLMs)中能否跨来源、模态和架构泛化?为解决此问题,我们构建了CMES(跨模态情感刺激),这是一个多来源的情感条件故事、真实面部表情、合成肖像和合成情感诱发场景的集合。对于每个刺激来源,我们从三个VLM中的每一个提取一组独立的六个Ekman情感向量。我们报告以下四个主要发现:1)图像衍生的情感向量形成与文本衍生向量相似的低维几何结构。效价在来源间相对稳定,而唤醒度的变化更大。2)文本和图像衍生的情感向量具有适度的余弦相似性,但仍显示出保留的跨模态对应关系。文本衍生的向量也能引导图像解释。3)即使原生余弦接近零,跨架构对应关系仍然存在。从通用ImageNet激活估计的变换恢复了对应关系和因果迁移,而无需使用六个情感向量或其标签。4)在对齐跨架构的表征后,我们构建了一个共享的情感子空间,该子空间保留了情感几何结构和选择性引导效应。相应的共识情感向量也能泛化到保留的第四种架构,涵盖两种模型规模。这些结果表明,情感表征可以跨来源、模态和架构共享关系结构和因果效应,即使单个向量方向不同。
英文摘要
Recent studies suggest that large language models encode emotion concepts as structured internal representations, but most existing work focuses on text and a single architecture. Therefore, we ask, do emotion concepts generalize across sources, modalities, and architectures in vision--language models (VLMs)? To address this, we construct CMES (Cross-Modal Emotion Stimuli), a multi-source collection of emotion-conditioned stories, real facial expressions, synthetic portraits, and synthetic emotion-evoking scenes. For each stimulus source, we extract a separate set of six Ekman emotion vectors from each of three VLMs. We report four main findings as follows: 1) Image-derived emotion vectors form a low-dimensional geometry similar to that of text-derived vectors. Valence is relatively stable across sources, while arousal varies more. 2) Text- and image-derived emotion vectors have modest cosine similarity but still show held-out cross-modal correspondence. Text-derived vectors can also steer image interpretation. 3) Cross-architecture correspondence remains even when native cosine is near zero. Transformations estimated from generic ImageNet activations recover both correspondence and causal transfer without using the six emotion vectors or their labels. 4) After aligning representations across architectures, we construct a shared emotion subspace that preserves affective geometry and selective steering effects. The corresponding consensus emotion vectors also generalize to a held-out fourth architecture at two model sizes. These results suggest that emotion representations can share relational structure and causal effects across sources, modalities, and architectures, even when individual vector directions differ.
发表机构
- Lappeenranta-Lahti University of Technology LUT(拉彭兰塔-拉赫蒂理工大学)
- Shanghai Jiao Tong University(上海交通大学)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- University of Oulu(奥卢大学)
- ELLIS Institute Finland(芬兰ELLIS研究所)
- Brno University of Technology(布尔诺理工大学)
机构由 AI 辅助整理,请以论文原文为准。