发表机构
School of Engineering, The University of British Columbia, Okanagan Campus(不列颠哥伦比亚大学工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对二进制删除信道,提出轨道约简与优化游程分布(ORD)方法,通过精确计算和神经-精确混合程序获得有限块长容量,并利用连续传递扩展到长块长,优于传统游程长度分布。
AI 中文摘要
对于在固定长度输入上运行的二进制删除信道,相关的性能指标是有限块长容量 $C_N(d)=\max_{p(x^N)}\mathrm{I}(X^N;Y)$,而不仅仅是无限块长极限 $C(d)$。我们证明,最优输入可以在补集和置换等价轨道上取常数,从而将优化简化为每个轨道一个权重,并引入了优化游程分布(ORD),这是一个 $N$ 参数的游程计数模型,对于 $N\le 3$ 与 $C_N(d)$ 一致,并且在所有游程计数恒定输入中是最优的。一种精确的嵌入计数动态规划评估这些结构化定律的 $\mathrm{I}$。一种混合神经-精确程序恢复了 $N\le 8$ 的认证 ORD 权重;变分批评器(InfoNCE、NWJ、DV/MINE、SMILE)仅用作内部搜索目标。直接得分函数学习 ORD 权重在 $N\ge 32$ 时坍缩为平坦的游程长度分布(RLD)。因此,我们引入了 ORD 连续传递:从精确的小 $N$ ORD 解中学习的归一化游程计数轮廓在目标长度上重新采样,直至 $N=512$。在中等删除概率下,传递的 ORD 始终优于 RLD;在 $d=0.1$ 时,增益大约在 $N=128$--$192$ 处消失。通过 Fertonani-Duman 的长度-熵不等式转换的精确 ORD 速率是 $C(d)$ 的有效下界;大 $N$ 输入的嵌套蒙特卡洛评估作为诊断报告,不声称是容量下界。一个固定 $N=100$ 的样本预算研究量化了嵌套 MC 偏差。所有主要报告的速率都是 $\mathrm{I}(X^N;Y)/N$ 的值。
英文摘要
In DNA data storage, racetrack memories, and packet networks, data are written as many short strands of fixed length. The receiver knows where each strand begins and ends, and deletions occur only inside a strand. The relevant limit for such systems is the block capacity $C_N(d)$, the largest mutual information between a length-$N$ input and the output of a binary deletion channel with deletion probability $d$. Because the boundaries are known, $C_N(d)/N$ is never smaller than the classical capacity $C(d)$. Computing $C_N(d)$ is hard because the input takes $2^N$ values. For $0\le d<1$, we show that the optimal input is unique, gives positive probability to every string, and is unchanged by complementing or reversing the strings. We then introduce the optimized run distribution (ORD), which assigns one probability to each number of runs; it is optimal for $N\le3$ and is solved with a numerical optimality certificate up to $N=16$. For longer strands, an exact recursion for the output distribution gives unbiased rate estimates, and an empirical Bernstein inequality turns them into confidence intervals. Within the statistical precision, a one-parameter Markov input performs as well as the best inputs found. For strands of 100-200 bits, known boundaries increase the rate by up to 0.026 bits per symbol beyond the best certified upper bound on the capacity of the unsegmented channel. Via Fano's inequality, the block capacity also bounds the rate of any code that uses a single strand; in the tabulated case, this bound is tighter than the best known finite-length bound for strands longer than about 90 bits. Conversely, block rates yield lower bounds on $C(d)$; at $d=0.1$, the bound is within $1.1\times10^{-4}$ bits/use of the best certified lower bound. Finally, InfoNCE estimates, even with the optimal critic, empirically lose most of their ability to rank inputs as the strand length grows.
Comments13 pages, 4 figures, 9 tables. v2: corrected orbit reduction (complement-reversal theorem); ORD values recomputed with numerical certificates; nested Monte Carlo replaced by exact-marginal estimator with confidence intervals, new Fano single-strand converse and InfoNCE analysis