叙事散文中的句法模式与文体功能:一种基于规则与机器学习的方法
Syntactic Patterns and Stylistic Functions in Narrative Prose: A Rule-Based and Machine-Learning Approach
AI总结:
本文提出一种基于规则与机器学习的方法,利用依存句法分析语料库中的线性化三元组模式,对叙事散文句子进行五类文体分类,最佳模型宏F1达0.948。
AI中文摘要:
本文提出了一项小规模定量实验,将叙事散文中的句法结构与文体功能联系起来。基于一个包含3,300个句子的依存句法分析语料库,我们通过一种透明的基于规则的程序,检查词元、通用词性标注和句法关系,为每个句子推导出五个类别——描述性、内省性、因果性、意识形态性和中性——的句子级文体标签。对于每个句子,我们构建其句法轮廓的紧凑表示,作为结合词元、词性标注和依存关系的线性化三元组序列。这些模式作为输入,输入到训练用于预测句子级文体的标准机器学习分类器中。表现最佳的模型在10折交叉验证下实现了0.948的宏F1分数。该实验完全使用Python和开源工具实现。我们的目标不是提出一个完整的文体理论,而是提供一个可复现且可扩展的工作流程,用于探索语法结构如何促进叙事解读。
英文摘要:
This paper presents a small-scale quantitative experiment that links syntactic structure to stylistic functions in narrative prose. Starting from a dependency-parsed corpus of 3,300 sentences, we derive sentence-level stylistic labels across five categories --- descriptive, introspective, causal, ideological, and neutral --- using a transparent rule-based procedure that inspects lemmas, universal part-of-speech tags, and syntactic relations. For each sentence we construct a compact representation of its syntactic profile as a sequence of linearised triples combining lemma, POS tag, and dependency relation. These patterns serve as input to standard machine-learning classifiers trained to predict sentence-level style. The best-performing model achieves a macro-F1 of 0.948 under 10-fold cross-validation. The experiment is implemented entirely in Python using open-source tools. Our goal is not to propose a fully fledged stylistic theory, but to offer a reproducible and extensible workflow for exploring how grammatical structure contributes to narrative interpretation.