arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

互信息约束的Chernoff瓶颈

Mutual Information Constrained Chernoff Bottleneck

Dier Tang, Guangyue Han

arXiv 2609.37994首次发表:更新:

发表机构

The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出互信息约束的Chernoff瓶颈,通过最大化Chernoff信息优化编码器,证明最优值特性并给出交替算法,在20 Newsgroups数据上以17%熵保留90%误差指数。

AI 中文摘要

经典信息瓶颈(IB)通过$I(U;Y)$衡量表示$U$(关于$X$)对目标$Y$的相关性,但这并不直接表征下游决策的误差。对于从多个独立编码的观测中推断出的二元假设$Y$,最优误差指数是给定$Y$时$U$的两个条件分布之间的Chernoff信息。我们研究互信息约束的Chernoff瓶颈,即寻求在速率约束$I(U;X) \leq R$下最大化该Chernoff信息的编码器。我们证明其最优值$C(R)$在$R = H(V)$之前严格递增,其中$V$合并了$X$中似然比相等的符号,超过该点后保持未压缩的指数,并且与IB曲线不同,$C(R)$不必是凹的。我们进一步证明$k+1$个输出足以达到$C(R)$,其中$k$是$V$的基数。我们提出一种交替算法,通过广义Blahut-Arimoto算法更新编码器,并通过非线性方程更新Chernoff参数$s$,并证明其迭代保持可行性,且Chernoff信息非递减并收敛。数值实验证实了理论,在20 Newsgroups语料库的真实主题检测数据上,将每个单词压缩到其熵的仅$17\\%$,保留了$90\\%$的误差指数,并几乎保持了未压缩分类器的准确性。

英文摘要

The classical information bottleneck (IB) measures the relevance of a representation $U$ of $X$ to a target $Y$ by $I(U;Y)$, which does not directly characterize the error of downstream decisions. For a binary hypothesis $Y$ inferred from many separately encoded observations, the optimal error exponent is the Chernoff information between the two conditional distributions of $U$ given $Y$. We study the mutual information constrained Chernoff bottleneck, which seeks an encoder that maximizes this Chernoff information subject to a rate constraint $I(U;X) \leq R$. We show that its optimal value $C(R)$ increases strictly up to $R = H(V)$, where $V$ merges the symbols of $X$ with equal likelihood ratio, remains at the uncompressed exponent beyond, and, unlike the IB curve, need not be concave. We further show that $k+1$ outputs suffice to attain $C(R)$, where $k$ is the cardinality of $V$. We propose an alternating algorithm that updates the encoder via a generalized Blahut--Arimoto algorithm and the Chernoff parameter $s$ via a nonlinear equation, and prove that its iterates remain feasible, with nondecreasing and convergent Chernoff information. Numerical experiments confirm the theory, and on real topic-detection data from the 20 Newsgroups corpus, compressing each word to only $17\%$ of its entropy retains $90\%$ of the error exponent and nearly the accuracy of the uncompressed classifier.

Comments31 pages, 3 figures, 2 tables. Feedback and comments are welcome

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑