arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19914cs.DScs.DB

ZigZag Trie:一种用于上下文查询的新型索引

ZigZag Trie: A Novel Index for Contextual Queries

Ling Li, Daniel Gibney, Sharma V. Thankachan, Rahul Shah, Grigorios Loukides, Solon P. Pissis

首次发表
浏览论文内容

中文总结 AI 辅助

针对文本中字符串上下文查询效率低的问题,提出ZigZag Trie全文索引,以线性时间空间构建,支持四种上下文查询,速度远超基线。

中文摘要 AI 辅助

对于在较长文本$T$中字符串$P$的上下文(即给定$q$,所有满足$|L|=|R|=q$且字符串$LPR$出现在$T$中的字符串对$(L,R)$的集合)的查询,人们越来越感兴趣。这种上下文查询在多个领域中都很重要,但高效地回答它们具有挑战性。这是因为应用中$T$的长度巨大,而现有索引并不直接编码给定$P$的上下文,而这对于高效地回答检索查询至关重要。我们的工作引入了ZigZag Trie(ZZT),一种专门解决这些挑战的新型全文索引。该索引重新组织文本,使得对于任何$P$,所有围绕$P$对称增长的字符串$L$和$R$都被分组到索引的一个公共子树中,从而允许高效地检索它们。我们展示了如何构建$T$的ZZT,其大小为$\nmathcal{O}(n)$(其中$n=|T|$),构建时间为$\nmathcal{O}(n\log n)$,空间为$\nmathcal{O}(n)$。在ZZT之上,我们设计了专门的索引,对于查询模式$P$,可以回答四种新型上下文查询:(I)找到在$T$中出现至少$\tau$次的最长字符串$LPR$,其中$\tau$固定;(II)找到在文本集合的至少$\tau$个文本中出现的最长字符串$LPR$,其中$\tau$固定;(III)报告$P$在$T$中不同上下文的总数;(IV)对于给定的$q$,根据给定的评分函数,检索$P$的得分最高的$k$个对$(L,R)$。我们的索引以最优时间回答类型I、II和III的查询,以近最优时间回答类型IV的查询。此外,给定ZZT,它们的大小、构建空间和构建时间在$n$上是线性或近线性的。使用真实的十亿字母数据集,我们表明我们的索引回答查询的速度比基线快几个数量级,并且在索引大小和构建空间及时间上表现相似或更好。

英文摘要

There is increasing interest in queries about the context of a string $P$ in a longer text $T$, i.e., the set of all string pairs $(L,R)$, with $|L|=|R|=q$, for a given $q$, such that the string $LPR$ occurs in $T$. Such contextual queries are important in several domains but are challenging to answer efficiently. This is because the length of $T$ in applications is massive and existing indexes do not directly encode the context of a given $P$, which is key for answering retrieval queries efficiently. Our work introduces the ZigZag Trie (ZZT), a new full-text index to specifically address these challenges. This index reorganizes the text so that, for any $P$, all possible strings $L$ and $R$ growing symmetrically around $P$ are grouped into a common subtree of the index, allowing their efficient retrieval. We show how to construct the ZZT of $T$, which has size $\mathcal{O}(n)$ where $n=|T|$, in $\mathcal{O}(n\log n)$ time and $\mathcal{O}(n)$ space. On top of ZZT, we design specialized indexes that, for a query pattern $P$, answer four new types of contextual queries: (I) finding the longest string $LPR$ that occurs at least $τ$ times in $T$, for a fixed $τ$; (II) finding the longest string $LPR$ that occurs in at least $τ$ texts of a text collection, for a fixed $τ$; (III) reporting the total number of distinct contexts of $P$ in $T$; and (IV) retrieving, for a given $q$, the $k$ pairs $(L,R)$ of $P$ with the highest scores according to a given scoring function. Our indexes answer queries of type I, II, and III in optimal time, and of type IV in near-optimal time. Moreover, their size, construction space, and construction time are linear or near-linear in $n$, given ZZT. Using real billion-letter datasets, we show that our indexes answer queries orders of magnitude faster than baselines and perform similarly or better in index size and construction space and time.

发表机构

  • King’s College London(伦敦国王学院)
  • University of Texas at Dallas(德克萨斯大学达拉斯分校)
  • North Carolina State University(北卡罗来纳州立大学)
  • Louisiana State University(路易斯安那州立大学)
  • The Cyprus Institute(塞浦路斯研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑