语素文字中的齐普夫缩写定律:汉字笔画数的编码理论界限
Zipf's Law of Abbreviation in a Logographic Script: Coding-Theoretic Bounds on Chinese Character Stroke Counts
浏览论文内容
中文总结 AI 辅助
本研究以汉字为对象验证齐普夫缩写定律,结合笔画数据库与语料库计算其最优性得分,发现汉字压缩上限独立于文字类型,简化汉字提升了最优性。
中文摘要 AI 辅助
齐普夫缩写定律(即高频形式倾向于更简短的趋势)是语言中最受支持的规律之一,近期研究已从验证该定律转向量化词典相对于原则性基准的压缩程度,该研究项目迄今已涉及表音文字和音节文字的词长。本研究将其拓展至语素文字,以笔画为发音成本单位,以汉字为编码形式。结合覆盖CJK基本区块全部20902个汉字的笔画顺序数据库,以及两个独立语料库(分别为2.589亿和1.933亿词次),研究发现:平均字类的笔画数为12.71,而文本中平均字次的笔画数仅为7.22。采用Petrini等人(2026)的双归一化最优性得分,简化字库的Omega值为0.668,复制语料库的Omega值为0.609,处于该作者报告的20种语言8种文字词长对应的62%-67%区间内及略低于该区间,表明压缩上限在很大程度上独立于文字类型和成本单位。语素文字还可计算绝对编码界限,因为笔画来自封闭的五元分类:精确的5元赫夫曼最优值为4.34笔,熵界限为4.28笔,因此观测系统比最优编码高1.66倍。该差距并非松弛而是结构:频率列表的5^(-l_i)的Kraft和为2.05,完整字库的为5.03,因此笔画串在一维上不可唯一解码;汉字通过笔画的二维排列而非序列来区分,放弃的压缩换来成分透明性。最后,将20世纪中期的简化改革视为受控压缩事件,发现其使最优性从0.555提升至0.668,节省集中于最常用的1000个汉字。
英文摘要
Zipf's law of abbreviation -- the tendency of frequent forms to be short -- is one of the best-supported regularities in language, and recent work has moved from demonstrating it to measuring how far lexicons are compressed relative to principled baselines. That programme has so far addressed word lengths in alphabetic and syllabic scripts. We transfer it to a logographic script, taking the stroke as the unit of articulatory cost and the Chinese character as the coded form. Combining a stroke-order database covering all 20,902 characters of the CJK basic block with two independent frequency corpora (258.9M and 193.3M tokens), we find that the mean character type costs 12.71 strokes but the mean character token in running text only 7.22. Using the dually normalised optimality score of Petrini et al. (2026), the simplified inventory reaches Omega = 0.668, with the replication corpus at 0.609 -- inside and just below the 62-67% band those authors report for word lengths across 20 languages and 8 scripts, suggesting a compression ceiling largely independent of script type and cost unit. A logographic script also makes absolute coding bounds computable, since strokes come from a closed five-element taxonomy: the exact 5-ary Huffman optimum is 4.34 strokes and the entropy bound 4.28, so the observed system is 1.66x above optimal coding. This gap is not slack but structure. The Kraft sum of 5^(-l_i) is 2.05 on the frequency list and 5.03 on the full inventory, so stroke strings are provably not uniquely decodable in one dimension; characters are disambiguated by the two-dimensional arrangement of strokes, not their sequence, and the forgone compression buys componential transparency. Finally, treating the mid-twentieth-century simplification reform as a controlled compression event, we find it raised optimality from 0.555 to 0.668, with savings concentrated in the 1,000 commonest characters.