arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19228cs.ITcs.FLmath.IT

组装理论与最小语法问题

Assembly Theory and the Smallest Grammar Problem

Wawrzyniec Bieniawski

首次发表
浏览论文内容

中文总结 AI 辅助

该研究评估8种压缩算法作为ASI近似器的有效性,提出Re-Pair T-NDR变体,发现增大字母表大小可延迟CAs偏离最优组装路径的发散点,确立压缩算法可作为ASI估计的实用工具。

中文摘要 AI 辅助

组装理论(Assembly Theory, AT)通过组装指数(Assembly Index, ASI)来量化复杂性,该指标被证明等价于最小直线程序(straight-line program, SLP)的大小。尽管这种等价性将AT与数据压缩领域联系起来,而数据压缩是一个历经数十年研究的领域,但特定算法作为ASI近似器的实际有效性在很大程度上仍未被探索。本研究在包含408个字符串的数据集上,对8种压缩算法进行了实证评估,这些算法涵盖基于语法的方案和字典方案(CAs),数据集包括368个合成字符串(含最大复杂性字符串以及由香农熵量化的符号分布平衡程度各异的字符串)和40个自然生物序列(基因组和蛋白质组)。我们提出了Re-Pair T-NDR,这是一种Re-Pair的分支定界冲突解决变体,它比本研究中研究的其他CAs提供了更紧的ASI上界。增大字母表大小会扩大CAs与ASI紧密匹配的范围,延迟CAs偏离最优组装路径的“发散点”。这些结果确立了压缩算法作为ASI估计实用工具的可行性。

英文摘要

Assembly theory (AT) quantifies complexity through the assembly index (ASI) -- a metric proven equivalent to the size of the smallest straight-line program (SLP). While this equivalence links AT to data compression -- a field shaped by decades of research -- the practical efficacy of specific algorithms as ASI approximators remains largely unexplored. This study provides an empirical evaluation of eight compression algorithms, spanning grammar-based and dictionary schemes (CAs) across a dataset of 408 strings -- comprising 368 synthetic strings (including max-complexity strings and strings with varying levels of symbol distribution balance, quantified by Shannon entropy) and 40 natural biological sequences (genomic and proteomic). We introduce Re-Pair T-NDR, a branch-and-bound tie-resolving variant of Re-Pair, providing a tighter upper bound on the ASI than the other CAs researched in this study. Increasing alphabet size extends the regime in which CAs closely track the ASI, delaying a ``divergence point'', where a CA deviates from the optimal assembly path. The results establish compression algorithms as practical tools for ASI estimation.

↑