发表机构
Independent Researcher Phoenix AZ USA; Independent Researcher(; )
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出自监督多模态嵌入框架,在无显式标签的情况下发现AI的审美结构,探讨其与人类情感标签的差异,可应用于RAG媒体组织与自动标注等场景。
AI 中文摘要
审美是艺术作品象征意义的重要组成部分,尽管具有主观性,人类仍会根据艺术唤起的情感对其进行分类,且不受模态限制。目前尚未充分探索的是,AI模型如何在没有显式标签或跨模态监督的情况下,对人类创作的媒介形成自身的审美分类。我们提出一种自监督框架,将文本、音频、图像和视频四种模态投影到共享的256维嵌入空间,并应用迭代聚类来发现审美结构。我们在一个弱监督多模态数据集上,讨论AI生成的聚类分配与人类情感 register 标签之间的差异。这项工作可应用于理解AI如何构建跨模态相似性、为检索增强生成(RAG)组织异构媒体集合以及自动数据标注。
英文摘要
Aesthetics are an important part of the symbolism of artistic works. Although subjective, humans categorize art based on the emotion evoked regardless of modality. What remains under-explored is how AI models form their own aesthetic categorization of human-produced media without explicit labels or cross-modal supervision. We present a self-supervised framework that projects four modalities (text, audio, image and video) into a shared 256-dimensional embedding space and applies iterative clustering to discover aesthetic structure. We discuss the divergence between AI-generated cluster assignments and human affective register labels on a weakly supervised multimodal dataset. This work has applications in understanding how AI structures cross-modal similarity, organizing heterogeneous media collections for Retrieval-Augmented Generation (RAG), and automated data labeling.
CommentsAccepted at ICMI Companion '26 (Companion Publication of the 28th ACM International Conference on Multimodal Interaction), October 5-9, 2026, Napoli, Italy. 4 figures, 2 tables