论文字破译中所用统计测度的非特异性
On the Non-Specificity of Statistical Measures Used in Script Decipherment
浏览论文内容
中文总结 AI 辅助
研究以生成性符号系统SIGIL为工具,发现文字破译中所用的统计测度不特定于语言,无法仅凭自身确立未破译符号系统编码言语。
中文摘要 AI 辅助
统计规律性常被用作证据,表明未破译的符号系统编码了语言,印度河文字的争论是典型例子。此类推断均依赖特异性:所报告的结果必须在合理的结构化非语言中显得异常。我们用专门构建的生成性符号系统SIGIL对这一前提进行建设性检验,其3000文本的核心语料库带有明确的组合意义,尽管没有符号具有音系值。评估前编制的文献注册表记录了54种方法,当可复现已发表的印度河文字结果和源定义的决策规则时,该方法可实现精确评分。通过这种方式评分的每一项标准,SIGIL都获得了与印度河语料库相同的类别,涵盖重复、方向不对称和词汇分布测试。熵、频率、位置、预测、分类器和网络测度的已声明重构也复现了熟悉的类印度河特征。随后的序列破译压力测试在同一语料库上对英语、梵语和泰米尔语达到了高词典覆盖率,而分组保留集的下降和不稳定密钥则揭示了该覆盖率的识别能力有多弱。该构建未确定印度河符号编码了什么:它表明所评估的测度检测到了组织性,却不特定于语言,因此无法仅凭自身确立编码的言语。
英文摘要
Statistical regularities are routinely offered as evidence that undeciphered sign systems encode language; the Indus script debate is the canonical example. Any such inference rests on specificity: the reported outcome must be unusual among plausible structured non-languages. We test that premise constructively with SIGIL, a purpose-built generative emblem system whose 3,000-text core corpus carries explicit compositional meanings although no sign has a phonological value. A literature registry compiled in advance of evaluation records 54 methods and admits a method to exact scoring when both the published Indus outcome and a source-defined decision rule can be reproduced. SIGIL receives the same category as the Indus corpus on every criterion scored this way, across repetition, directional-asymmetry, and lexical-distribution tests. Declared reconstructions of entropy, frequency, positional, predictive, classifier, and network measures reproduce the familiar Indus-like signatures as well. A sequential decipherment stress test then reaches high dictionary coverage for English, Sanskrit, and Tamil on the same corpus, while grouped held-out declines and unstable keys reveal how little that coverage identifies. The construction does not decide what the Indus signs encode: it shows that the evaluated measures detect organization without being specific to language, and therefore cannot, on their own, establish encoded speech.