ChronoLens:测量跨时间、语言和语言层面的语言变化
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
浏览论文内容
中文总结 AI 辅助
ChronoLens框架结合多语言模型等技术,分析1803-2026年五议会传统的海量文本,揭示同语言内各语言层面变化幅度相当、不同语言变化轨迹有差异,为跨语言历史语言变化研究提供方法。
中文摘要 AI 辅助
历史语言变化会影响形态、句法、语义和语用层面,但计算研究通常采用不兼容的表示方法来考察这些层面,因此无法确定它们是否跨语言共同演化。我们通过探究变化的幅度和方向如何在语言层面、语言以及单一分析空间内的历史时期之间变化来解决这一问题。我们引入ChronoLens框架,该框架结合了冻结的多语言语言模型、特征对齐跨编码器和事后语言干预,并将其应用于1803年至2026年间五个议会传统的4498万份文档和约172亿个token。所得稀疏表示与语言统计数据的一致性远强于密集嵌入或池化稀疏自编码器(相关系数ρ=0.72,对比0.29和0.28),且揭示同一语言内的形态、句法、语义和语用通常变化幅度相当,而不同语言在变化的时间、程度和方向上存在显著差异。这些发现表明,历史语言变化是一个结构化的多维过程:相似的幅度可能掩盖不同的轨迹,而有意义的跨语言比较需要同时测量距离和方向。
英文摘要
Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages. We address this problem by asking how the magnitude and direction of change vary across linguistic levels, languages, and historical periods within a single analytical space. We introduce ChronoLens, a framework that combines frozen multilingual language models, feature-aligned crosscoders, and post-hoc linguistic interventions, and apply it to 44.98 million documents and approximately 17.2 billion tokens from five parliamentary traditions spanning 1803--2026. The resulting sparse representations agree substantially more strongly with linguistic statistics than dense embeddings or a pooled sparse autoencoder ($ρ=0.72$ versus $0.29$ and $0.28$), and reveal that morphology, syntax, semantics, and pragmatics generally change by comparable amounts within a language, while languages differ markedly in when, how far, and in which direction they change. These findings show that historical language change is a structured, multidimensional process: similar magnitudes can conceal different trajectories, and meaningful cross-linguistic comparison requires measuring both distance and direction.