arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25184cs.CL

人类语言中上下文持久性的标度律

A scaling law of contextual persistence in human language

Elan Barenholtz

首次发表
浏览论文内容

中文总结 AI 辅助

研究人类语言中单词顺序排列的规律,用大语言模型测量上下文持久性函数P(d),发现其在多个语料库中按1/d衰减,建立了人类语言中上下文持久性的标度律。

中文摘要 AI 辅助

人类语言在单词层面(频率、词汇增长)和词对层面(跨距离共现)呈现出规律结构。本文表明,单词的顺序排列——意义的核心决定因素——也遵循类似规律。使用大语言模型作为概率探测器,我们测量了在距离d处,与相同单词打乱顺序相比,先前上下文所带来的目标困惑度的降低。这种差异,即上下文持久性函数P(d),分离了排列的影响。在跨越六个语系以及书面和口语模式的十个语料库中,P(d)大约按1/d衰减。在打乱顺序和合成控制中该效应消失,在独立探测器中得到复制,且在领域原生模型下的基因组或蛋白质序列中未出现。接近1的指数在对数时间尺度上大致均匀地分布上下文影响。结果建立了人类语言中上下文持久性的标度律。

英文摘要

Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we show that the arrangement of words in sequence -- a central determinant of meaning -- obeys a comparable law. Using large language models as probabilistic probes, we measured the reduction in target perplexity conferred by prior context at distance d beyond that of the same words scrambled; this difference, the contextual persistence function P(d), isolates the influence of arrangement. Across ten corpora spanning six language families and written and spoken modalities, P(d) decayed approximately as 1/d ($P(d) \propto d^{-α}$, mean $α= 1.04$; median $r^2 = 0.96$). The effect vanished in scrambled and synthetic controls, replicated across independent probes, and did not appear in genomic or protein sequences under domain-native models. An exponent near 1 distributes contextual influence approximately uniformly across logarithmic timescales. The results establish a scaling law of contextual persistence in human language.

发表机构

  • Florida Atlantic University(佛罗里达大西洋大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑