arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

熵最小化选择中的任意放置问题及残差熵公式

The Arbitrary-Placement Problem in Entropy-Minimizing Selection, and a Residual-Entropy Formulation

Alyssa H. Shin, Claire H. Shin

arXiv 2610.05925首次发表:更新:

发表机构

California Institute of Technology; Cornell University(加州理工学院; 康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对熵最小化选择中的任意放置问题,提出残差熵公式,证明其安全性与排序条件,并在路由、类不平衡及视频预测任务中验证有效性。

AI 中文摘要

基于熵的选择目标存在一个根本性的退化问题:最小化香农熵 $H(p_A)$ 会奖励自信的选择,无论所选候选是否具有信息量。我们通过残差熵 $D = H(p_A) - H(p_\beta)$ 来解决这一局限性,其中 $p_\beta$ 由候选信任权重诱导。我们证明了精确恒等式 $D = -\mathrm{KL}(p_A\Vert p_\beta) - \Delta$,其中 $\Delta$ 衡量分数诱导分布与信任分布是否偏好相同的候选。边界情况确立了基本安全性:在均匀信任下,$D\leq0$ 自动成立,因此平等信任、非饥饿状态永远不会受到惩罚,而在任何 one-hot 极限下,$D\to0$ 无论选择哪个候选。对于发生选择的中间状态,我们证明当候选按信任排序与按信息量排序逐对一致时,$D\leq0$,并基于领先候选相对于其竞争者的优势推导出更紧的证书。这些结果独立于候选评分函数,并适用于静态和动态变化的信息。使用基于梯度的混合专家路由器进行的实验证实,排序条件可以在实际优化过程中成立,并表明当候选不可互换且选择直接使用而非平均时,正确的排序能提高下游性能。在路由之外,基于边际的重新加权在类不平衡任务中匹配或优于固定强度基线,而在生产视频预测系统中,信息性选择将 MSE 降低了约 20\\%,并迁移到相关物种。因此,残差熵为选择提供了安全标准,并为决定选择何时具有信息量提供了可用信号。

英文摘要

Entropy-based selection objectives suffer from a fundamental degeneracy: minimizing Shannon entropy $H(p_A)$ rewards confident selection regardless of whether the selected candidate is informative. We address this limitation with the residual entropy $D = H(p_A) - H(p_β)$, where $p_β$ is induced by candidate trust weights. We prove the exact identity $D = -\mathrm{KL}(p_A\Vert p_β) - Δ$, where $Δ$ measures whether the score-induced distribution and trust profile favor the same candidates. Boundary cases establish basic safety: under uniform trust, $D\leq0$ automatically, so an equal-trust, non-starving state is never penalized, while at any one-hot limit, $D\to0$ regardless of the selected candidate. For the intermediate regime where selection occurs, we prove that $D\leq0$ when candidate ordering by trust agrees pairwise with ordering by informativeness, and derive a tighter certificate based on the leading candidate's margin over its competitors. These results are independent of the candidate-scoring function and apply to both stationary and dynamically changing information. Experiments with a gradient-based mixture-of-experts router confirm that the ordering conditions can hold during real optimization and show that correct ordering improves downstream performance when candidates are non-interchangeable and selections are used directly rather than averaged. Beyond routing, margin-based reweighting matches or outperforms fixed-strength baselines in a class-imbalance task, while informative selection in a production video-prediction system reduces MSE by approximately 20$\%$ and transfers to a related species. Residual entropy, therefore, provides a safety criterion for selection and a usable signal for deciding when that selection is informative.

Comments28 pages, including references and appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑