arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Flesch-Kincaid可读性仅取决于主题模型下长文本的主题分布

Flesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic Models

Yo Ehara

arXiv 2608.23327首次发表:更新:

发表机构

Tokyo Gakugei University(东京学艺大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究基于带句子边界标记的主题模型,发现长文本的Flesch-Kincaid可读性仅取决于主题分布,通过实验在Brown和BNC语料库中验证了主题向量对FKGL的高预测性。

AI 中文摘要

Flesch阅读易读性(FRE)和Flesch-Kincaid年级水平(FKGL)是广泛使用的英文可读性评分,由相同的两个文档统计量计算得出,但它们在长文档上的稳定性并不意味着对词汇构成的不变性。令人惊讶的是,在带有显式句子边界标记的主题模型下,这两种评分几乎必然收敛为仅通过两个标量速率依赖于文档主题分布的确定性函数:在长文本极限下,所有评分变化均由主题构成而非任何剩余可读性信号介导。该理论覆盖了这两个公式,而实验评估了FKGL。在秩[1, q, s]=3的固定混合模型中,穿过内部主题向量的纤维是局部(K-3)维的,而规则等评分水平集是局部(K-2)维且弯曲的。在两个平衡语料库Brown和书面BNC的折外评估中,从一个文档一半的内容词推断出的主题向量预测另一半的FKGL,相关系数r分别为0.779和0.884。在Brown语料库中,将主题预测与体裁和平均内容词音节数相加后,ΔR²=0.002,置信区间包含零;在BNC语料库中,对应的折半增量为0.024,在五次K=100拟合中有四次为正(中位数为0.021)。由于推断出的主题可能也吸收了体裁、语域和风格的信息,因此不将这些结果解释为关于人类可读性或因果效应的证据。

英文摘要

Flesch Reading Ease (FRE) and the Flesch-Kincaid Grade Level (FKGL) are widely used readability scores for English computed from the same two document statistics, yet their stability on long documents need not imply invariance to lexical composition. Surprisingly, under a topic model with an explicit sentence-boundary token, both scores converge almost surely to deterministic functions of the document topic distribution through just two scalar rates: in the long-text limit, all score variation is mediated by topical composition rather than any residual readability signal. The theory covers both formulae, while the experiments evaluate FKGL. In a fixed admixture with rank[1, q, s] = 3, fibres through interior topic vectors are locally (K-3)-dimensional, whereas regular iso-score level sets are locally (K-2)-dimensional and curved. In out-of-fold evaluation on two balanced corpora, Brown and the written BNC, a topic vector inferred from one document half's content words predicts the other half's FKGL at r = 0.779 and 0.884, respectively. On Brown, adding the topic prediction to genre and mean content-word syllable count yields $ΔR^2$ = 0.002, with a confidence interval spanning zero; on the BNC, the corresponding split-half increment is 0.024, positive in four of five K = 100 fits (median 0.021). Because inferred topics may also absorb genre, register, and style, we do not interpret these results as evidence about human readability or causal effects.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑