arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27783cs.ITmath.IT

自适应编码中训练数据的最优加权

Optimal Weighting of Training Data in Adaptive Coding

Yuriy Reznik

首次发表
浏览论文内容

中文总结 AI 辅助

针对自适应编码中训练数据与源不匹配的问题,提出通过缩放KT估计器计数的最优加权,推导出冗余最小化权重,并给出三种实现方式,实验表明可显著降低冗余。

中文摘要 AI 辅助

自适应编码器通常以训练数据作为初始输入,这些数据被转化为上下文表或随附的字典。训练数据通常与正在编码的源不匹配,编码器必须决定对其信任的程度。我们将这种信任作为设计变量:一个在$m$元字母表上的Krichevsky--Trofimov(KT)估计器,其训练计数按$\xi\in[0,1]$缩放。对于长度为$\ell$的训练序列,其与消息源的Kullback--Leibler散度为$D$ nats每符号,则使冗余最小化的权重为$\xi^*=d/(2\ell D+d)$,其中$d=m-1$。有效训练长度$\xi^*\ell$遵循调和规律:任何不匹配都将可用训练信息限制在$d/(2D)$个符号以内,无论收集了多少数据。三种可实现的选取器提供未知的$D$:离线方式,根据训练集的分布;通过插件循环,根据解码前缀;以及通过双重通用混合,无需任何估计。在Calgary和Canterbury语料库的十个英文文本上,加权方法消除了经典KT码(即$\xi=0$端点)冗余的至多$22.7\\%$,并且在每个文件上都优于两个端点。提供了开源实现。

英文摘要

Adaptive coders are commonly primed with training data, turned into context tables or shipped dictionaries. The training data generally do not match the source being encoded, and the coder must decide how much to trust them. We make that trust a design variable: a Krichevsky--Trofimov (KT) estimator over an $m$-ary alphabet whose training counts are scaled by $ξ\in[0,1]$. For a training sequence of length $\ell$ at Kullback--Leibler divergence $D$ nats per symbol from the message source, the redundancy-minimizing weight is $ξ^*=d/(2\ell D+d)$, $d=m-1$. The effective training length $ξ^*\ell$ follows a harmonic law: any mismatch caps the usable training information at $d/(2D)$ symbols, however much was collected. Three implementable selectors supply the unknown $D$: offline, from the spread of the training set; by a plug-in loop, from the decoded prefix; and by a twice-universal mixture, with no estimation at all. On the ten English texts of the Calgary and Canterbury corpora, weighting removes up to $22.7\%$ of the redundancy of the classical KT code, the $ξ=0$ endpoint, and beats both endpoints on every file. An open-source implementation is provided.

发表机构

  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑