传统与神经方法在多语言可读性评估中的分析
Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment
浏览论文内容
中文总结 AI 辅助
本研究通过SHAP和TCAV方法,分析多语言Transformer模型是否内化了传统可读性评估中的语言特征,发现其恢复表面、句法和词汇多样性信号,但特征对齐因模型和语言而异。
中文摘要 AI 辅助
基于Transformer的模型在自动可读性评估(ARA)中表现出色,但基于特征的模型仍在使用,因为它们的预测与语言属性相关联。这一点很重要,因为可读性标签是主观的且依赖于评分者,因此在高噪声真实标签上的高准确率可能反映的是表面模式,而非定义难度的语言结构。我们使用ReadMe++数据集测试了Transformer是否在阿拉伯语、英语、法语、印地语和俄语中内化了与传统模型相同的特征。Shapley加性解释(SHAP)识别出驱动传统分类器的特征,我们将其用作TCAV概念集来探测多语言XLM-R和特定语言编码器。Transformer恢复了表面长度、句法和词汇多样性信号,并反映了传统模型的序数CEFR结构。对齐程度因模型家族、语言和层而异,特定语言编码器比XLM-R更清晰地跟踪传统模型。高线性可分性并不总是意味着方向性影响,这限制了线性探针用于基于计数的可读性特征。
英文摘要
Transformer-based models excel at Automatic Readability Assessment (ARA), yet feature-based models remain in active use because their predictions tie back to linguistic properties. This matters because readability labels are subjective and rater-dependent, so high accuracy on noisy ground truth may reflect surface patterns rather than the linguistic structure that defines difficulty. We test whether transformers internalize the same features as traditional models across Arabic, English, French, Hindi, and Russian using the ReadMe++ dataset. Shapley Additive Explanations (SHAP) identify the features driving traditional classifiers, which we then use as TCAV concept sets to probe multilingual XLM-R and language-specific encoders. Transformers recover surface-length, syntactic, and lexical-diversity signals, and reflect the ordinal CEFR structure of the traditional models. Alignment varies by model family, language, and layer, with language-specific encoders tracking traditional models more clearly than XLM-R. High linear separability does not always imply directional influence, limiting linear probing for count-based readability features.
发表机构
- Harvard University(哈佛大学)
- Massachusetts Institute of Technology(麻省理工学院)
- Kensho Technologies
机构由 AI 辅助整理,请以论文原文为准。