语言模型生成的小说呈现出压缩的形式多样性
Novels generated by language models show compressed formal variation
- University of Maryland(马里兰大学)
- University of West Bohemia(西波希米亚大学)
- Charles University(查理大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究对比不同模型生成及人类撰写的小说语料库,发现AI生成小说的形式多样性被压缩,句子结构等指标的变异性远低于人类小说,且存在方差与相关性过度封闭现象。
AI中文摘要:
大型语言模型能够生成完整的小说,但关于其多次生成内容的形式多样性水平,目前信息甚少。本研究未询问能否识别出单个段落为AI生成,而是探究重复的AI生成能否产生人类语料库中存在的相同多样性范围。本文对比了六个基于生成来源和目标风格的语料库:20篇采用GPT-5.5 Thinking生成的19世纪英国现实主义风格小说、20篇采用Qwen3-14B生成的19世纪英国现实主义风格小说、20篇采用上述每种模型生成的当代零风格小说、205部19世纪英国人类撰写的小说,以及65部当代人类撰写的零风格小说。在文档层面,研究采用了MATTR-500、香农熵、平均句子长度、可读性和标点符号率等测量指标。最稳健可靠的结果是句子结构的压缩:重复生成的小说在句子结构上的差异远小于人类小说;可读性、标点符号和小说内部句子长度变异性的测量指标也存在压缩现象;词汇测量指标同样呈压缩状态,仅Qwen零风格的MATTR例外。尽管GPT和Qwen具有不同的平均风格特征,但它们缺乏稳定的跨测量相关性模式。因此,本文区分了方差过度封闭(代表小说间有限的形式范围)和更具体的相关性过度封闭现象,即单篇AI生成小说可能在风格上类似人类小说,但AI生成小说的整体集合占据的形式范围要窄得多。
英文摘要:
While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather than asking whether individual passages can be identified as AI-generated, this study asks whether repeated AI generation can produce the same range of diversity which is found across human corpora. This paper contrasts six corpora based on generation source and target style: twenty novels generated using GPT-5.5 Thinking in a nineteenth-century British realist style, twenty novels generated using Qwen3-14B in a nineteenth-century British realist style, twenty novels generated using each of these models in a contemporary zero style, 205 nineteenth-century human-written British novels, and sixty-five contemporary human-written Zero-Style novels. At the document level, the research includes MATTR-500, Shannon entropy, average sentence length, readability, and punctuation rate measurements. The most robust and reliable result is compression of sentence structure. Repeated generations produce novels that vary far less from one another in sentence structure than human novels do. Compression is also present in the measures of readability, punctuation, and sentence length variability within novels. Lexical measures tend to be similarly compressed, with the exception of Qwen Zero-Style MATTR. Despite having distinct mean stylistic profiles, GPT and Qwen lack a stable pattern of cross-measure correlation. This article therefore distinguishes between variance overclosure, which represents a limited formal range between novels, and a more specific phenomenon of correlational overclosure. This means that an individual AI-generated novel may resemble human fiction stylistically, while a collection of AI-generated novels occupies a much narrower formal range.