arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

glyph不是字母,token不是词,空格不是空格:伏尼契手稿的单位究竟不是什么

A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not

Liudmila Rozanova, Alexander Temerev

arXiv 2608.17096首次发表:更新:

发表机构

University of Geneva; International Institute for Applied Systems Analysis (IIASA)(日内瓦大学; 国际应用系统分析研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究检验伏尼契手稿的三个常见假设,发现其glyph非字母、token非词、空白非词空格,揭示其单位特征,现有模仿模型无法复现其关键特性,强调需基于测量而非假设解释手稿。

AI 中文摘要

伏尼契手稿(Beinecke MS 408)通常基于三个未明确说明的假设进行分析:其 glyph 是字母、空白之间的字符串是词、每个空白都是词空格。我们使用Zandbergen-Landini转写文本,结合匹配的散文、密码和伪文本对照,以及抄本层级的重采样,对这三个假设进行检验。结果显示三个假设均不成立,且失败具有共同特征:伏尼契手稿的顺序性体现在token的边缘以及token之间的渐变边界,而非token本身的序列中。glyph的规律性过强,无法对应任何经检验的明文的一对一替换(条件熵为2.7比特,而拉丁语、意大利语和英语的条件熵约为3.5比特),而是分解为抄本层级稳定的重复多符号单位。token可构成合理的词汇,但一个token的身份对下一个token的预测仅占token熵的不到1%,低于所有匹配对照(2%-10%);而token边缘的glyph共享0.2比特的互信息,高于任何散文对照。空白分为两类:转录者标记为不确定的分隔符表现类似词内部的衔接,在页面上物理宽度更窄(独立图像坐标得出AUC为0.905,小型盲法墨水审计也得到相同结果),且即使在学习前擦除所有空格,仍会被学习到的单位跨越。该特征也是区分性依据。已发表的伏尼契模仿密码及自引用文本生成器均重现了低熵、单位尺度、弱token顺序及校准替换攻击的零结果,但均未重现边缘glyph耦合或开放的、稀有词占比高的词汇(70%为单例类型,对照为41%及59%-60%)。因此,对手稿的任何解释都必须基于测量而非假设,建立从glyph、token和分隔符到字母、词和词空格的对应关系,而这些测量正是实现该对应关系的依据。

英文摘要

The Voynich manuscript (Beinecke MS 408) is usually analysed on three unstated assumptions: that its glyphs are letters, that the strings between blanks are words, and that every blank is a word space. We test all three against the Zandbergen-Landini transliteration with matched prose, cipher, and pseudo-text controls and quire-level resampling. None holds, and the failures share a shape: the order in Voynichese sits at the edges of tokens and at graded boundaries between them, not in the succession of tokens themselves. Glyph regularity is too strong for one-to-one substitution of any tested plaintext (conditional entropy 2.7 bits against about 3.5 for Latin, Italian, and English) and resolves instead onto a quire-stable scale of recurrent multi-symbol units. Tokens form a plausible vocabulary, yet the identity of one token predicts the next by under 1% of token entropy, below every matched control (2-10%), while the glyphs at token edges share 0.2 bits of mutual information, more than in any prose control. Blanks fall into two regimes: the separators transcribers marked uncertain behave like word-internal junctures, are physically narrower on the page (AUC 0.905 from independent image coordinates, with the same sign in a small blind ink audit), and are crossed by learned units even when every space is erased before learning. This profile is also what discriminates. A published Voynich-imitating cipher and a self-citation text generator both reproduce the low entropy, the unit scale, the weak token order, and the null result of a calibrated substitution attack; neither reproduces the edge-glyph coupling or the open, hapax-rich vocabulary (70% singleton types against 41% and 59-60%). Any account of the manuscript must therefore earn, rather than assume, the step from glyphs, tokens, and separators to letters, words, and word spaces, and these are the measurements on which to do so.

Comments33 pages, 7 figures, 3 appendices. Analysis code and data are included as ancillary files and mirrored at https://github.com/lrozanova/voynich-units

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑