arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19493cs.ITmath.COmath.IT

通过子串计数的线性哈希实现多项式更大的删除码

Polynomially larger deletion codes by linear hashing of substring counts

Eyal En Gad

AI总结:

本文提出基于子串计数线性哈希的删除码构造,将纠正两次删除的冗余度上界从4log2n改进为3log2n,并推广到t次删除,通过标签类提取和二染色处理混淆词。

AI中文摘要:

我们证明,长度为 $n$ 的二进制码在纠正两次删除时,冗余度为 $3\log_2n+O(\log_2\log_2n)$。此前最佳上界的前导系数为 $4$,自1965年以来未变,而已知最佳下界的系数为 $2$。更一般地,纠正 $t\ge2$ 次删除的码存在冗余度为 $(2t-1)\log_2n+O_t(\log_2\log_2n)$,改进了系数 $2t$。我们从子串计数的随机线性哈希的一个标签类中提取码,其标签数比直接构造少约 $n$ 倍。仍共享标签的可混淆词通过二染色分离,在丢弃具有奇环的分量中的词之后,这些词很少,因为奇环迫使沿其编辑重叠。

英文摘要:

We show that binary codes of length $n$ correcting two deletions exist with redundancy $3\log_2n+O(\log_2\log_2n)$. The previous best upper bound had leading coefficient $4$, unchanged since 1965, while the best known lower bound has coefficient $2$. More generally, codes correcting $t\ge2$ deletions exist with redundancy $(2t-1)\log_2n+O_t(\log_2\log_2n)$, improving the coefficient $2t$. We extract a code from one label class of a random linear hash of substring counts, with about $n$ times fewer labels than a direct construction. Confusable words that still share a label are separated by a two-colouring after discarding the words in components with odd cycles, and these are few because an odd cycle forces the edits along it to overlap.

↑