发表机构
Universitat Pompeu Fabra; Dokuz Eylul University; McLean Hospital; Harvard Medical School; University of Zurich; ETH Zurich; Institut Català de Recerca i Estudis Avançats (ICREA)(庞培法布拉大学; 多库兹·埃吕尔大学; 麦克莱恩医院; 哈佛医学院; 苏黎世大学; 苏黎世联邦理工学院; 加泰罗尼亚研究与高级研究院(ICREA))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用大型语言模型发现精神病患者连贯言语中信息压缩存在缺陷,表现为惊异度差异减弱和内在维度降低,且与语法组织相关。
AI 中文摘要
在语言任务上表现接近人类的大型语言模型(LLMs)已经改变了神经多样性条件下语言的研究。LLMs以高维向量(嵌入)的形式提供语言输入的表示,并从这些嵌入计算下一个词元的预测。先前的跨语言证据表明,在精神病中,LLM表示的内在维度(ID)较低,平均惊异度(预测误差)较高,这两者共同表现为一种复杂度降低。我们假设这些指标反映了精神病中信息压缩的普遍缺陷,该缺陷与语法组织相关,而语法组织正是实现预测的基础。我们将惊异度差异操作化为基于词频估计的惊异度与基于上下文语言模型(对语法组织敏感,超越词汇概念)的惊异度之间的差值。使用一个包含144名土耳其语使用者的数据集,其中包括106名精神分裂症谱系障碍(SSD)患者——56名慢性精神分裂症(SZH)、33名首发精神病(FEP)和17名分裂情感性障碍(SZA)——以及38名健康对照者。我们报告:(1)所有临床组的惊异度差异相对于对照组均减弱,且与词数无关;(2)SZH和FEP的压缩性(内在维度)降低;(3)句法复杂性和压缩性均能预测惊异度差异。这些结果进一步细化了先前证实的精神病中语义空间几何形状的改变,表明该障碍存在更广泛的信息压缩缺陷,其机制基础在于语法的运算。
英文摘要
Large language models (LLMs) with human-like performance on linguistic tasks have transformed the study of language in neurodiverse conditions. LLMs provide representations of linguistic input in the form of high-dimensional vectors (embeddings), and next-token predictions computed from these embeddings. Previous crosslinguistic evidence suggests a complexity reduction in the form of both lower intrinsic dimensionality (ID) of LLM representations and higher mean surprisal (prediction error) in psychosis. We hypothesized that these metrics reflect a general deficit of information compression in psychosis, linked to grammatical organization as what enables predictions in language.We operationalized surprisal difference as the difference between surprisal as estimated from word frequency and surprisal as based on a contextual LM, which is sensitive to grammatical organization over and above lexical concepts. Using a dataset of 144 Turkish speakers, including 106 patients with schizophrenia-spectrum disorders (SSD) - 56 with chronic schizophrenia (SZH), 33 with first-episode psychosis (FEP), and 17 with schizoaffective disorder (SZA) - and 38 healthy controls. We report: (1) Surprisal difference is attenuated in all clinical groups relative to controls, independently of word count; (2) Compressibility (intrinsic dimension) is reduced in SZH and FEP; (3) Syntactic complexity and compressibility both predict surprisal difference. These results, further refining an alteration in the geometry of the semantic space in psychosis as previously attested, suggest a broader deficit in information compression in this disorder, with a mechanistic underpinning in the operations of grammar.