基于五种大语言模型的英文歌词文化分析重复测量研究
A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models
查看机构详情
- School of Applied and Creative Computing, Purdue University(普渡大学应用与创意计算学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究通过五种大语言模型对英文歌词的四种社会建构进行重复测量,发现其可靠性因建构而异,自尊的可靠性最强,且共识标签含可学习信号,需报告稳定性与收敛性。
中文摘要 AI 辅助
大语言模型(LLMs)正越来越多地被用于对文化文本进行标注,其规模是人工编码人员难以实现的。然而,在将这些模型的输出视为潜在社会建构的测量结果之前,必须确定这些测量是否可靠。本研究评估了五种大语言模型作为零样本标注器,对英文歌词中表达的四种社会建构:自尊、自我控制、寻求归属感和寻求认可进行标注的情况。通过对大型歌词语料库的重复标注,我们考察了基于大语言模型的测量的三个属性:重复运行的一致性、模型间的收敛性以及共识标签向监督分类的可迁移性。研究结果表明,基于大语言模型的测量在不同建构间并非都可靠。自尊在不同模型间表现出最强的重复测量可靠性,而寻求认可通常稳定性较差;自我控制和寻求归属感表现出中等但依赖于模型的可靠性。下游分类进一步表明,大语言模型的共识标签包含可学习的信号,尽管可迁移性本身并不能确立建构效度。因此,在将大语言模型标注视为文化分析中的可扩展测量之前,应报告重复测量稳定性和跨模型收敛性。
英文摘要
Large language models (LLMs) are increasingly used to annotate cultural texts at scales that are impractical for human coders. However, before their outputs are treated as measurements of latent social constructs, it is necessary to establish whether those measurements are reliable. This study evaluates five LLMs as zero-shot annotators of four social constructs expressed in English song lyrics: self-esteem, self-control, seeking belonging, and seeking recognition. Using repeated annotations of a large lyric corpus, we examine three properties of LLM-based measurement: consistency across repeated runs, convergence across models, and transferability of consensus labels to supervised classification. The findings show that LLM-based measurement is not uniformly reliable across constructs. Self-esteem exhibits the strongest repeated-measurement reliability across models, while seeking recognition is generally less stable; self-control and seeking belonging show intermediate but model-dependent reliability. Downstream classification further indicates that consensus LLM labels contain learnable signal, although transferability does not itself establish construct validity. Repeated-measurement stability and cross-model convergence should therefore be reported before LLM annotations are treated as scalable measurements in cultural analytics.