SAGE:一个用于评估叙事中解释性文学质量的分层框架
SAGE: A Hierarchical Framework for Evaluating Interpretive Literary Quality in Narratives
浏览论文内容
中文总结 AI 辅助
SAGE是一个六层评估框架,用于衡量叙事中的解释性文学质量,发现大语言模型在情感表征上接近人类,但在文化批评和哲学深度上差距显著。
中文摘要 AI 辅助
评估叙事的文学质量需要评估解释性维度(文化表征、情感深度和哲学参与),而现有的自然语言生成指标无法衡量这些维度。我们引入了SAGE,一个六层评估框架,将基于规则的可观察文本属性评估与基于大语言模型的解释性质量评估(这些质量源自文化理论、情感理论和存在主义哲学)分离开来。每个解释性层通过多轮迭代的大语言模型评估并辅以独立的交叉验证进行评估,实现了跨评估模型稳定的测量级可靠性(98.8%的一致性,>94%的评估者间一致性)。在100篇短篇故事的600次评估中验证后,我们的核心发现是一个系统性的能力边界:情感-心理表征接近人类水平,而文化批评和哲学深度的差距约为其两倍。大语言模型生成的叙事在所有三个层面上的得分甚至低于商业类型小说。我们将此解释为可模式复制的文学能力(可从训练语料库中学习)与需要文化定位和哲学参与的立场要求能力之间的边界,而仅靠模式匹配无法提供后者。
英文摘要
Assessing the literary quality of narratives requires evaluating interpretive dimensions (cultural representation, emotional depth, and philosophical engagement) that existing NLG metrics cannot measure. We introduce SAGE, a six-layer evaluation framework that separates rule-based assessment of observable textual properties from LLM-based evaluation of interpretive qualities drawn from cultural theory, affect theory, and existentialist philosophy. Each interpretive layer is assessed through multi-round iterative LLM evaluation with independent cross-validation, achieving measurement-grade reliability (98.8% convergence, >94% inter-rater agreement) stable across evaluator models. Validated on 600 evaluations across 100 short stories, our central finding is a systematic capability boundary: emotional-psychological representation approaches human levels, while cultural critique and philosophical depth exhibit approximately double the gap. LLM-generated narratives score below even commercial genre fiction on all three layers. We interpret this as a boundary between pattern-reproducible literary capacities learnable from training corpora and stance-requiring ones demanding cultural positioning and philosophical engagement that pattern matching alone cannot provide.
发表机构
- Mercy University(梅西大学)
- IBM T.J. Watson Research Center(IBM TJ·沃森研究中心)
机构由 AI 辅助整理,请以论文原文为准。