AI诗歌的作者归属与美学评价:以俳句为例的案例研究
Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku
浏览论文内容
中文总结 AI 辅助
本研究通过少样本提示生成多类LLM的日本俳句,结合人类创作进行问卷测试,发现美学评价与作者识别分离,且随模型改进,人类辨别能力可能下降。
中文摘要 AI 辅助
本文研究了当代大型语言模型(LLMs)对日本俳句的生成与人类评价,聚焦于在受限诗歌形式中的作者感知与美学判断。采用少样本提示策略,在一组异构的大型语言模型上生成了日本俳句,这些模型包括开源和闭源系统、中等规模和大规模架构、具有原生或适配日语支持的模型,以及多语言专有模型。这些AI生成的俳句与人类创作的俳句混合后,通过问卷形式呈现给东京日本大学的学生。调查评估了受访者能否区分AI生成与人类创作的俳句,以及哪些线索影响了他们的判断。识别准确率因模型而异。GPT-5、Gemini 2.5和StableLM-7B的表现约为随机水平(约0.50),而LLM-JP、Gemma-2B和LLaMA-2显示出中等可检测性(约0.59-0.67)。然而,识别结果强烈依赖于具体条目。流畅性、连贯性、诗性及相关美学维度的评分预测了感知的人类性,但并未预测正确分类,这表明存在一种与美学评价相关的归因偏差,并揭示了美学评价与真实作者检测之间的分离。扩展分析还考察了生成约束的遵循情况、参与者层面的特征,以及基于LLM的探索性俳句作者评价。总体而言,研究结果表明,随着LLMs的改进,表面层面的创作合理性可能会降低人类在受限诗歌环境中可靠辨别的能力。
英文摘要
This paper investigates the generation and human evaluation of Japanese haiku by contemporary Large Language Models (LLMs), focusing on authorship perception and aesthetic judgment within a constrained poetic form. Using a few-shot prompting strategy, Japanese haiku were generated across a heterogeneous set of large language models, including open- and closed-source systems, medium-scale and large-scale architectures, models with native or adapted Japanese support, and multilingual proprietary models. These AI-generated haiku were combined with human-written ones and presented in a questionnaire distributed to students at Japanese universities in Tokyo. The survey assessed whether respondents could distinguish between AI-generated and human-written haiku and which cues informed their judgments. Recognition accuracy varied across models. GPT-5, Gemini 2.5, and StableLM-7B performed at approximately chance level (approx 0.50), whereas LLM-JP, Gemma-2B, and LLaMA-2 showed moderate detectability (approx 0.59-0.67). However, recognition was strongly item-dependent. Ratings of fluency, coherence, poeticness, and related aesthetic dimensions predicted perceived humanness but not correct classification, indicating an attribution bias linked to aesthetic evaluation and revealing a dissociation between aesthetic evaluation and true authorship detection. The extended analysis additionally examines generation-constraint adherence, participant-level characteristics, and exploratory LLM-based evaluations of haiku authorship. Overall, the findings suggest that as LLMs improve, surface-level creative plausibility may reduce reliable human discrimination within constrained poetic settings.
发表机构
- Sapienza University of Rome(罗马大学)
- Shibaura Institute of Technology(芝浦工业大学)
机构由 AI 辅助整理,请以论文原文为准。