一个好的预测器能错得多离谱?预测KL消失时的发散更新
How Wrong Can a Good Predictor Be? Diverging Updates with Vanishing Predictive KL
浏览论文内容
中文总结 AI 辅助
本研究证明在平稳对称高斯HMM中,即使内部更新差距无界,预测KL仍可消失,表明内部更新差距本身不必然导致预测失败,并隔离了其与预测成本间的两个缺失环节。
中文摘要 AI 辅助
准确的预测后验并不需要准确逼近贝叶斯更新。我们证明,在平稳对称高斯隐马尔可夫模型中,对于每个固定的有限$K\ge2$,更新映射之间的无界差距可以与消失的预测KL共存。精确贝叶斯混合和显式确定性径向滤波器作用于相同的$K-1$个信念坐标。当$q\to0^+$时,在最坏情况下,它们在中心化logit空间中的分离度至少随自然置信度尺度$L_K(q)$线性增长,而它们的分类$D_{\mathrm{KL}}(\mathrm{exact}\\|\mathrm{radial})$在相同的显式见证点上消失。沿着平稳HMM轨迹,滤波后验之间的期望终端KL也在$H(q)=\lceil-\log(q)/c\rceil+1$处收敛到零。没有切换的典型块将两个滤波器驱动到一个共同的置信锥中,其中softmax曲率抑制了它们的分歧;单个高斯最大事件控制自适应噪声。在$K\in\{2,4,8\}$上使用等间距高斯分布的扫描说明了相反的趋势,长时域上的二元控制比较了饱和与非饱和递归。该结果隔离了内部更新差距与预测成本之间的两个缺失环节:分离状态对期望损失的贡献以及解码器敏感性。因此,即使是无界的内部更新差距本身也不能证明预测失败。该构造在$K$上固定,不提供压缩何时无害的通用准则,也不表征内部差距何时必然导致任务损失。
英文摘要
Accurate posterior prediction need not require accurate approximation of Bayesian updates. We prove that an unbounded gap between the update maps can coexist with vanishing predictive KL for every fixed finite $K\ge2$ in a stationary symmetric Gaussian HMM. Exact Bayesian mixing and an explicit deterministic radial filter act on the same $K-1$ belief coordinates. As $q\to0^+$, their separation in centered logits in the worst case grows at least linearly in the natural confidence scale $L_K(q)$, while their categorical $D_{\mathrm{KL}}(\mathrm{exact}\|\mathrm{radial})$ vanishes at the same explicit witness. Along stationary HMM trajectories, the expected terminal KL between filtered posteriors also converges to zero at $H(q)=\lceil-\log(q)/c\rceil+1$. Typical blocks without switches drive both filters into a common confidence cone, where softmax curvature suppresses their disagreement; a single Gaussian maximal event controls adaptive noise. A sweep with equally spaced Gaussians over $K\in\{2,4,8\}$ illustrates the opposing trends, and binary controls at long horizons compare saturating and nonsaturating recurrences. The result isolates two missing links between internal update gaps and predictive cost: the contribution of separating states to expected loss and decoder sensitivity. Thus even an unbounded internal update gap does not by itself certify predictive failure. The construction is fixed in $K$ and does not provide a universal criterion for when compression is harmless or characterize when internal gaps must incur task loss.
发表机构
- Boston University(波士顿大学)
- Shanghai Jiao Tong University(上海交通大学)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。