AI 中文总结
研究针对数字对象跨模态情感比较难题,提出CD-MED方法,通过各模态情感识别模型输出转化为共享描述符,在共同情感空间表示对象,实现跨域情感比较、检索、推荐及可视化。
AI 中文摘要
数字对象通过不同模态表达情感,如电影含视觉场景、音频等,歌曲有旋律、歌词等。现有情感识别模型多针对特定模态,难以直接比较此类对象。本文提出CD-MED,一种跨域多模态情感描述符,可在共同情感空间表示异构数字对象。各模态经自身情感识别模型处理,输出转化为共享描述符,既保留各模态信息,又呈现对象综合情感概况。还在效价-唤醒空间可视化,实现跨域情感比较、检索、推荐及可视化。
英文摘要
Digital objects express emotions through different modalities. For example, a movie may include visual scenes, audio, dialogue, and facial expressions, while a song may contain melody, rhythm, lyrics, and vocal tone. Because existing emotion recognition models are usually modality-specific, it is difficult to compare such objects directly. This paper proposes CD-MED, a Cross-Domain Multimodal Emotion Descriptor for representing heterogeneous digital objects in a common emotional space. Each modality can be processed by its own emotion recognition model, and the resulting emotional outputs are transformed into a shared descriptor. The descriptor preserves information from individual modalities while also allowing an integrated emotional profile of the object. For interpretation, CD-MED is visualized in the valence-arousal space: position represents affective coordinates, color denotes emotion category, size indicates intensity, and shape shows the modality. This unified representation enables emotion-based comparison, retrieval, recommendation, and visualization across different domains such as movies, songs, images, and books.
CommentsSubmitted to 2026 Joint 14th International Conference on Soft Computing and Intelligent Systems and 27th International Symposium on Advanced Intelligent Systems (SCIS&ISIS 2026)