AI 中文总结
研究大语言模型生成文本与人类文本差异,通过分析n元语法统计分布及定性分析,揭示大语言模型风格不足,得出风格和语义问题不可完全分离的结论。
AI 中文摘要
先前关于大语言模型生成文本的研究已表明其在定量和定性方面与人类创作的文本存在差异。大语言模型生成的文本在风格上与人类写作不同,有独特的文本“感觉”,且语义范围比人类受限。本文指出大语言模型生成文本中n元语法统计分布的简单且一致的模式,通过定性分析揭示了大语言模型风格的不足。由于高阶n元语法与语义内容相关,得出风格和语义问题并非完全可分离的结论。
英文摘要
Prior work on LLM-generated text has demonstrated quantitative and qualitative departures from text produced by humans. LLM-generated texts differ from human writing in style, resulting in a characteristic textual "feel," while the semantic range of LLMs is much restricted compared to that of humans. In this contribution, I note simple but consistent patterns in the statistical distribution of n-grams within LLM-generated text. Via qualitative analysis of these n-grams, I reveal deficiencies in LLM style. Because higher-order n-grams correlate to semantic content, I conclude that questions of style and semantics are not cleanly separable.