AI生成音乐中味觉-声音对应关系的跨文化评估
Cross-cultural evaluation of taste-sound correspondences in AI-generated music
浏览论文内容
中文总结 AI 辅助
本研究通过阿根廷、意大利、日本三国实验,评估AI生成音乐中味觉-声音对应关系的跨文化差异,发现差异源于量表使用与感知结构,需区分反应风格偏差与真实感知重组。
中文摘要 AI 辅助
声音调味研究表明,听众会将系统性的味觉与情感意义赋予声音,而文本到音乐的生成式人工智能近来被用于将味觉提示转化为音乐刺激。此类模型习得的味觉-声音对应关系是否能在其验证的文化语境之外成立,仍未得到检验。我们将一项单国研究扩展至在阿根廷、意大利和日本开展的三国在线实验,共招募361名参与者。参与者首先对从四个味觉提示(甜、酸、苦、咸)生成的基础模型与微调后MusicGen片段进行偏好选择,随后对微调后片段在12个味觉、情感和热觉描述符上进行评分。阿根廷和意大利的参与者偏好微调后模型,但日本参与者无此偏好;且在所有三个群体中,咸的提示对应的对应关系最弱。各国的评分存在显著差异,但在对参与者内部评分进行标准化后,国家的主效应不再可检测,而提示与描述符映射的交互作用基本保持不变。因此,大部分明显的跨文化差异可归因于量表使用的差异;不过仍存在结构性成分。此外,探索性因子分析显示,12个描述符在每个群体中沿不同的潜在维度组织。这些结果表明,AI介导的声音调味的跨文化差异在两个层面运作:一是将味觉赋予特定刺激的整体层面,二是这些赋予的关系结构层面。因此,在不同人群中评估生成式音乐系统时,应区分反应风格偏差与真正的感知重组。
英文摘要
Sonic seasoning research has shown that listeners attribute systematic gustatory and emotional meaning to sound, and text-to-music generative artificial intelligence has recently been used to render gustatory prompts as musical stimuli. Whether the taste-sound correspondences acquired by such models hold beyond the cultural context in which they were validated remains untested. We extended a single-country study to a three-country online experiment conducted in Argentina, Italy, and Japan (N = 361). Participants first indicated their preference between base and fine-tuned MusicGen excerpts generated from four taste prompts (sweet, sour, bitter, salty), and then rated fine-tuned excerpts on twelve taste, emotion, and thermal descriptors. Preference for the fine-tuned model was confirmed in Argentina and Italy but not in Japan, and the salty prompt yielded the weakest correspondence in all three cohorts. Ratings differed substantially between countries, yet the main effect of country was no longer detectable once ratings had been standardized within participant, whereas the interactions characterizing the mapping of prompts onto descriptors remained essentially unchanged. Much of the apparent cross-cultural divergence is therefore attributable to differences in scale use; a structural component nevertheless persists. In addition an exploratory factor analysis indicated that the twelve descriptors were organized along different latent dimensions in each cohort. These results indicate that cross-cultural variation in AI-mediated sonic seasoning operates at two levels: the overall level at which taste is attributed to a given stimulus, and the relational structure of those attributions. Evaluations of generative music systems across populations should accordingly distinguish response-style bias from genuine perceptual reorganization.
发表机构
- University of Padova(帕多瓦大学)
- University of Trento(特伦托大学)
- University of Gastronomic Sciences(美食科学大学)
- Universidad Nacional de Tres de Febrero(三月十五日国立大学)
- Ritsumeikan University(立命馆大学)
机构由 AI 辅助整理,请以论文原文为准。