自然数的可变长度格雷码
Variable-length Gray codes for the Natural Numbers
AI总结:
研究提出一种自然数的可变长度格雷码\(\V\),通过特定构造方式获得。证明了其具有双射、编辑距离为\(1\)、码字长度单调非递减等性质,实现了\(\Bits\)的哈密顿枚举,还讨论了其压缩行为及在特定编码家族中的应用。
AI中文摘要:
反射二进制格雷码用于排列整数的固定长度二进制表示,使得连续数字仅在一位上不同。但其有用性受固定字长\(b\)限制,既限制了可表示数字范围为\(2^{b}-1\),又在小整数上浪费位。我们引入可变长度格雷码\(\V\),它是从自然数到所有有限二进制字符串集的全函数,通过对\(n + 1\)取反射格雷码、丢弃前导零并删除单个前导一来获得。我们证明了此构造的四个属性:一是\(\V\)是自然数与\(\Bits\)之间的双射,是完整码;二是连续整数码字的莱文斯坦(编辑)距离总是\(1\);三是码字长度随编码数字单调非递减且等于\(\lfloor \log_2 (n + 1) \rfloor\);四是从\((k - 1)\)位到\(k\)位码字的转换恰好在\(n = \sum_{i = 0}^{k} 2^{i} = 2^{k + 1} - 1\)处。该码在单位编辑步骤下实现了\(\Bits\)的哈密顿枚举,同时保持了近乎最优的自适应长度分布。我们讨论了它相对于固定长度格雷码的压缩行为及其在图、超图、分子和符号回归表达式的\(\emph{Isal*}\)指令集字符串编码家族中的应用,其中编码之间的莱文斯坦距离用作结构相似性代理。
英文摘要:
The modular $b$-ary Gray code arranges fixed-length $b$-ary representations of intervals of natural numbers so that consecutive numbers differ in a single digit. Its usefulness, however, is tied to a fixed word length. We introduce, for every integer base $b\ge 2$, a \emph{variable-length Gray code} $V_{b}$: a bijection from the natural numbers onto the set of all finite strings over the $b$-symbol alphabet $\{0,1,\dots,b-1\}$. The construction orders the natural numbers by codeword length into blocks and lists each block along the modular $b$-ary Gray code of the within-block offset, with its leading digit incremented modulo $b$. We prove four properties for arbitrary $b$. First, $V_{b}$ is a bijection, a complete code that assigns exactly one codeword to every finite string over the alphabet, including the empty string. Second, the Levenshtein (edit) distance between the codewords of two consecutive integers is always one. Third, codeword length is monotone non-decreasing in the encoded number. Fourth, the transition to codewords of length $k$ occurs exactly at $n=N_{k}=\sum_{i=0}^{k-1}b^{i}$, where $V_{b}(n)=1\,0^{k-1}$. The presented code, therefore, realizes a Hamiltonian enumeration under unit edit steps while retaining a near-optimal, self-adapting length profile. These properties suggest a use in machine learning: large language models routinely emit natural numbers as symbol strings, yet the positional notations they rely on are neither complete -- strings with leading zeros are invalid or redundant -- nor locally stable, since incrementing a number may rewrite many symbols at once. Because $V_{b}$ is a complete code and moves by a single edit between consecutive integers, it removes both obstacles and is a natural candidate representation for numeric tokens; we develop this argument and discuss the code's compression behavior relative to fixed-length Gray codes.