AI 中文总结
研究通过证明组合引理,将其应用于字符串后缀树,提出算法框架,存储频繁子串计数,利用报告查询计算不频繁子串,实现高效索引,支持多种字符串计数应用,提升查询时间。
AI 中文摘要
我们证明了一个基本但强大的组合引理:在任何有 \(L\) 个叶子的根树中,深度小于其叶子后代数量的节点数量最多为 \(L\)。对于长度为 \(n\) 的任何字符串 \(T\),将此引理直接应用于 \(T\) 的后缀树可得,长度小于其在 \(T\) 中出现次数的 \(T\) 的子串数量最多为 \(n\)。这种组合见解通过以下算法框架为字符串计数问题带来了具有最佳查询时间的空间高效数据结构:在预处理步骤中存储 \(T\) 中最多 \(n\) 个“频繁”子串的计数,并使用报告查询来计算“不频繁”子串。我们的框架就像一个方便的黑匣子,将具有报告时间 \(\mathcal{O}(|P| + |\textsf{Occ}_T(P)|)\) 的索引提升为支持在时间 \(\mathcal{O}(|P|)\) 内进行计数查询,其中 \(P\) 是查询模式,\(\textsf{Occ}_T(P)\) 是 \(P\) 在 \(T\) 中的出现集。作为应用,我们展示了用于连续出现、加权序列、带效用的字符串和非重叠出现的高效索引。
英文摘要
We prove an elementary yet powerful combinatorial lemma: in any rooted tree with $L$ leaves, the number of nodes whose depth is smaller than the number of their leaf descendants is at most $L$. For any string $T$ of length $n$, a direct application of this lemma to the suffix trie of $T$ yields that the number of substrings of $T$ whose length is smaller than their number of occurrences in $T$ is at most $n$. This combinatorial insight leads to space-efficient data structures with optimal query times for string counting problems via the following algorithmic framework: store the counts for the at most $n$ ``frequent'' substrings of $T$ in a preprocessing step, and use a reporting query to count for the ``infrequent'' substrings. Our framework acts as a convenient black box, lifting indexes with reporting time $\mathcal{O}(|P|+|\textsf{Occ}_T(P)|)$ to support counting queries in time $\mathcal{O}(|P|)$, where $P$ is the queried pattern and $\textsf{Occ}_T(P)$ is the set of occurrences of $P$ in $T$. As applications, we show efficient indexes for consecutive occurrences, weighted sequences, strings with utilities, and non-overlapping occurrences.
CommentsESA 2026
DOI:10.4230/LIPIcs.ESA.2026.68