arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

惊奇值并非一种理论

surprisal is Not a Theory

Andrés Buxó-Lugo, Aniello De Santo, Morgan Grobol, Ryan J. Hubbard, Cassandra L. Jacobs

arXiv 2607.20208首次发表:更新:

AI 中文总结

研究指出惊奇值理论虽被视为计算层面解释,但大语言模型的发展未免除建模者相关表征决策,不加批判使用LLM - 惊奇值会模糊模型承诺,通过分析表明算法和架构选择影响语言模型概率计算,建议相关研究者重新评估概率互换做法。

AI 中文摘要

惊奇值理论常被视为一种计算层面的解释。我们认为,尽管计算层面的叙述被用于支持计算心理语言学中的“表征无关研究”,但大语言模型(LLMs)所体现的向黑箱系统的转变,并未使使用惊奇值度量的建模者免除计算层面表征所要求的表征决策。实际上,对LLM - 惊奇值的不加批判的使用模糊了不同模型的表征和算法层面的承诺。通过三项分析,表明算法和模型架构的选择在语言模型概率计算中起重要作用。建议希望测试惊奇值理论的研究人员重新评估将大语言模型概率视为可互换的做法。

英文摘要

Surprisal Theory is often characterized as a computational-level explanation per (Marr, 1982). We argue in this work that, even though a computational level narrative has been used to support "representation-agnostic research" within computational psycholinguistics, the movement toward black box systems embodied by large language models (LLMs) does not exempt modelers using the surprisal metric from the representational decisions required by computational-level characterizations. In fact, we argue that the uncritical use of LLM-surprisal obfuscates the representational and algorithmic-level commitments of different models. In three analyses, we show that the choice of algorithm and model architecture play significant roles in the computation of language model probabilities. We advise that researchers who wish to test Surprisal Theory re-evaluate the practice of treating large language model probabilities as interchangeable

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑