AI 中文总结
本文提出一种结合字符3-gram、UMAP和Burrows' Delta的归纳文体计量方法,用于恢复政治文本中的潜在作者结构,并在多语言、多模式语料库中验证其有效性。
AI 中文摘要
政治文本很少仅由名义上的发言者单独撰写。推文、演讲、报告和官方声明往往由工作人员起草、编辑或协调统一,然而政治科学对这些隐藏作者留下的文体痕迹关注有限。本文开发并压力测试了一种归纳文体计量学方法,用于恢复政治传播中的潜在作者结构,该方法结合了字符3-gram特征、UMAP降维和Burrows' Delta。我们将该方法应用于六个语料库,这些语料库在长度(从推文到长文档)、模式(书面和口头)以及语言(英语和匈牙利语)上各不相同。该方法在两种语言的正式法律散文中恢复了近乎不相交的分析者指纹,将一位政治家的推文分类为经过验证的子集,同时揭示了额外的见解,并区分了有脚本和即兴的演讲。然而,它未能解决有脚本语料库中的个别演讲稿撰写人。因此,基于频率的文体计量学是一个强大的工具,根据作者信号强度和机构编辑情况,可以揭示与立法研究、政治传播和政策研究相关的作者痕迹。
英文摘要
Political texts are rarely authored by the nominal speaker alone. Tweets, speeches, reports, and official statements are drafted, edited, or harmonized by staff, yet political science has paid limited attention to the stylistic traces these hidden authors leave behind. This paper develops and stress-tests an inductive stylometric approach for recovering latent authorship structure in political communication, combining character 3-gram features with UMAP dimensionality reduction, and Burrows' Delta. We apply the approach to six corpora that vary in length (from tweets to long documents), in mode (written and oral), and in language (English and Hungarian). The approach recovers near-disjoint analyst fingerprints in formal legal prose in both languages, sorts a politician's tweets into validated subsets while uncovering additional insights, and distinguishes scripted from improvised speech. It fails, however, to resolve individual speechwriters within scripted corpora. Frequency-based stylometry is thus a powerful tool that, depending on authorial signal strength and institutional editing, can uncover authorship traces relevant to legislative studies, political communication, and policy research.
Comments34 pages, 17 figures, 3 tables. Includes appendices A-D (validation materials and discriminant validity checks). Under review