arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向大字母集的分层单纯形架构

A Layered Simplex Architecture for Large Alphabets

Meir Feder, Yaniv Fogel, Ruediger Urbanke

arXiv 2608.19908首次发表:更新:

发表机构

School of Electrical and Computer Engineering; Tel Aviv University; School of Computer and Communication Sciences; EPFL(电气与计算机工程学院; 特拉维夫大学; 计算机与通信科学学院; 洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种分层单纯形架构的新贝叶斯估计器,其构造简单无需调参,在大字母集概率估计任务中性能优于Good-Turing等专业化方法,还可揭示数据、字母集大小与深度的缩放规律。

AI 中文摘要

在对数损失下对大字母集进行概率估计是一个已被广泛研究的问题,著名方法包括Good-Turing估计器。本文引入并研究一种具有四个显著特性的新贝叶斯估计器:第一,其构造极为简单:对概率单纯形上的独立均匀抽样逐元素相乘后再重新归一化;深度是唯一的结构参数,对深度取平均可消除调参需求。第二,该混合模型的遗憾(相对于已知信源码的额外码长)存在显式且可高效计算的表达式。第三,尽管结构简单且无调优常数,该估计器在多种合成与真实文本基准测试中,与包括Good-Turing在内的更专业化方法相比仍具竞争力。第四,其可计算的遗憾使我们能识别数据、字母集大小与深度的缩放规律:对于指数大于1的Zipf目标,当样本仅揭示字母集的一小部分时,遗憾可简单表述为与已发现符号集的描述长度密切匹配——每比特描述对应1比特代码,且每个符号需额外成本;数据指数即为新符号被发现的速率。

英文摘要

Probability estimation over large alphabets under log loss is a well-studied problem, with celebrated methods such as the Good-Turing estimator. We introduce and study a new Bayesian estimator with four notable properties. First, its construction is exceptionally simple: multiply independent uniform draws from the probability simplex coordinate-wise and renormalize. Depth is the only structural parameter, and averaging over depths eliminates the need to tune it. Second, the regret of the resulting mixture, the excess code length it pays relative to a code that knows the source, admits an explicit and efficiently computable expression. Third, despite its simplicity and lack of tuned constants, the estimator is competitive across a diverse set of synthetic and real-text benchmarks with substantially more specialized methods, including Good-Turing. Fourth, the tractability of its regret allows us to identify scaling laws in data, alphabet size, and depth. For Zipf targets with exponent above one, the regret has a simple reading as long as the sample reveals only a small fraction of the alphabet. It closely matches the description length of the set of discovered symbols, at one bit of code per bit of description, plus a further cost per symbol. The data exponent is therefore the rate at which new symbols are discovered.

Comments30 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑