AI 中文总结
研究变量到变量码压缩率的统计特性,通过推导精确矩公式、Edgeworth近似等,得出相关闭式公式,还将分析扩展到马尔可夫源,在编码理论上给出结构分解,如应用于Khodak码体现其改进性能在较小偏差常数上。
AI 中文摘要
变量到变量(V2V)长度码将源序列解析为可变长度的短语,并将每个短语映射为通常具有不同随机长度的二进制码字。编码n个短语后,实现的压缩率$R_n=\Lambda_n/\Sigma_n$(总码字长度除以总源符号数)是码的渐近率$\rho$的有限样本对应物,仅在$n\to\infty$时收敛。本文首先推导了给定离散无记忆源(DMS)时$R_n$所有整数矩的精确公式。具体而言,我们获得了每个矩$\E\{R_n^k\}$的闭式公式,它是仅涉及对$(L,\ell)$(源符号中的短语长度和比特中的码字长度)的单短语矩生成函数的一维积分。从这些矩中,我们推导了$R_n$累积分布函数(CDF)的Edgeworth近似,它比中心极限定理(CLT)近似更准确。使用拉普拉斯积分方法,我们还推导了偏差常数$C=\lim_{n\to\infty}n(\E\{R_n\}-\rho)$和方差常数$\lim_{n\to\infty}n\cdot\Var\{R_n\}$的显式闭式公式。该分析通过具有闭式冗余公式的状态索引矩阵扩展到马尔可夫源。在编码理论方面,我们将V2V长度码视为有限状态编码器,并应用广义Kraft不等式获得压缩率下限,并给出偏差系数的结构分解,该分解在可变到固定(V2F)长度码、固定到可变(F2V)长度码和V2V长度码之间清晰分离。应用于Bugeaud、Drmota和Szpankowski的Khodak码,这种分解表明其改进的性能体现在其较小的偏差常数上。
英文摘要
A variable-to-variable (V2V) length code parses a source sequence into phrases of variable length and maps each phrase to a binary codeword of, generally, a different random length. After encoding $n$ phrases, the realized compression ratio $R_n=Λ_n/Σ_n$ -- total codeword length over total source-symbol count -- is the finite-sample counterpart of the code's asymptotic rate $ρ$, to which it converges only as $n\to\infty$. This paper first derives exact formulas for all integer moments of $R_n$ for a given discrete memoryless source (DMS). Specifically, we obtain a closed-form formula for every moment $\E\{R_n^k\}$ as a one-dimensional integral involving only single-phrase moment generating functions of the pair $(L,\ell)$ -- the phrase length, in source symbols, and codeword length, in bits. From these moments we derive an Edgeworth approximation to the cumulative distribution function (CDF) of $R_n$ that is substantially more accurate than the central limit theorem (CLT) approximation. Using the Laplace method of integration, we also derive explicit closed-form formulas for the bias constant $C=\lim_{n\to\infty}n(\E\{R_n\}-ρ)$ and for the variance constant $\lim_{n\to\infty}n\cdot\Var\{R_n\}$. The analysis extends to Markov sources via state-indexed matrices with a redundancy formula obtained in closed form. On the coding-theoretic side, we cast V2V length codes as finite-state encoders and apply a generalized Kraft inequality for a compression-rate lower bound, and give a structural decomposition of the bias coefficient that separates cleanly across variable-to-fixed (V2F) length codes, fixed-to-variable (F2V) length codes, and V2V length codes. Applied to the Khodak code of Bugeaud, Drmota, and Szpankowski, this decomposition shows that its improved performance is reflected in its smaller bias constant.
Comments37 pages, 2 figures, submitted for publication