发表机构
Uppsala University(乌普萨拉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究通过构建含30篇文本的基准从大语言模型推理痕迹提取文学质量隐含理论,经五次复制实验得模型判断写作质量的相关结论;还通过对经典散文段落退化探究该理论,发现其判断整体、作者特定且对结构更敏感,对相关领域有启示。
AI 中文摘要
在文学研究和计算语言学中,写作“好”的标准一直是个持久的问题。我们进行了两项研究来探究具备推理能力的大语言模型如何评估文学质量。研究1构建了一个包含30篇真实文本的基准,跨越六个质量层级,从经典文学到匿名论坛帖子,从模型推理痕迹中提取其质量隐含理论。在五次对DeepSeek的复制实验中,模型平均层级分类准确率达79.3%。痕迹揭示了一个一致的既定理论:模型看重意图而非正确性,优先考虑技巧、深度和独特声音。一项对风格匹配但不可识别段落的熟悉度实验表明,来源识别可能会提高分数,不过这因经典原文与研究者创作的仿作之间的真实质量差异而混淆。研究2通过对五篇经典散文段落进行系统退化来探究该理论。我们应用了六种操作——词汇简化、节奏扁平化、意象去除、声音通用化、结构简化和综合退化,并重新评估每个版本。词汇简化导致的质量损失最小(0.41±0.46分),远低于结构(2.78分)或声音(2.34分)损失。综合退化具有毁灭性(-5.64分)但次可加性。与文心一言的探索性比较显示出相同的广泛定性模式。这些研究共同表明,大语言模型对写作质量的判断是整体的、作者特定的,并且对结构特征比对词汇特征更敏感,这对自动写作反馈和计算美学具有启示意义。
英文摘要
What makes writing "good" remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality. In Study 1, we construct a benchmark of 30 real texts spanning six quality tiers, from canonical literature to anonymous forum posts, and extract the model's implicit theory of quality from its reasoning traces. Across five DeepSeek replications, the model achieves 79.3% mean tier-classification accuracy. The traces reveal a consistent stated theory: the model values intentionality over correctness, prioritizing craft, depth, and distinctive voice. A familiarity experiment with style-matched but unrecognizable passages suggests that source recognition may inflate scores, although this is confounded by genuine quality differences between canonical originals and researcher-written pastiches. In Study 2, we probe this theory through systematic degradation of five canonical prose passages. We apply six manipulations - vocabulary simplification, rhythm flattening, imagery removal, voice genericization, structure simplification, and combined degradation - and reevaluate each version. Vocabulary simplification causes the smallest quality loss (0.41 +/- 0.46 points), far below structure (2.78) or voice (2.34) loss. Combined degradation is devastating (-5.64) but subadditive. An exploratory comparison with Qwen QwQ shows the same broad qualitative pattern. Together, these studies suggest that LLM judgments of writing quality are holistic, author-specific, and more sensitive to structural than lexical features, with implications for automated writing feedback and computational aesthetics.