arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语义正则表达式的精确两轮适应性和轮次层次结构

Sharp Two-Round Adaptivity and Round Hierarchies for Semantic Regular Expressions

Runzhou Li, Hongfei Fu, Qingkai Shi, Peisen Yao

arXiv 2607.22799首次发表:更新:

AI 中文总结

研究语义正则表达式中适应性及轮次层次结构,通过多项式大小电路等方法,确定了两轮适应性极值能力、最大非自适应与自适应比率等,展示了完整轮次层次结构,区分了语义信息获取等成本。

AI 中文摘要

语义正则表达式(SemREs)将外部布尔谓词附加到匹配跨度上,使预言调用的数量和顺序成为核心资源。对于固定表达式和单词,我们用多项式大小的单调跨度电路表示成员关系,并将最优语义评估与布尔决策树评估联系起来。我们渐近精确地确定了适应性的极值能力。对于每个$E\geq2$,存在一个语法大小为$\Theta(E)$的一元、无星、语义深度为一的实例,有$E$个基本预言键且只有单位长度语义跨度,其一轮成本为$E$,而其精确的两轮和无限制确定性成本为$\log_2 E+\frac{1}{2}\log_2\log_2 E+O(1)$。因此,最大的非自适应与自适应比率为$(1+o(1))E/\log_2E$。第二个受限族展示了完整轮次层次结构:其最优$R$轮成本为$\Theta(RE^{1/R})$。两种构造都允许在固定字母表$\{0,1,\#\}$上进行单谓词实现,语义跨度为对数长度,总表示大小为$O(E\log^2 E)$。在逐点误差$\delta<1/2$和最坏情况预期成本下,对于每个有$E$个基本键的实例,随机非自适应复杂度恰好为$(1 - 2\delta)E$。最后,对于固定单词$w$和$h$个谓词名,精确的随机极小极大值为$(1 - 2\delta)h\ sd(w)$,其中$sd(w)$计算不同子串值;跨度界限$s$用$sd_s(w)$代替$sd(w)$来计算。这些结果区分了语义信息获取、并行延迟和局部符号匹配成本。

英文摘要

Semantic regular expressions (SemREs) attach external Boolean predicates to matched spans, making both the number and the sequentiality of oracle calls central resources. For a fixed expression and word, we represent membership by a polynomial-size monotone span circuit and identify optimal semantic evaluation with Boolean decision-tree evaluation. We determine the extremal power of adaptivity asymptotically sharply. For every $E\ge2$, there is a unary, star-free, semantic-depth-one instance of syntax size $Θ(E)$ with $E$ essential oracle keys and only unit-length semantic spans whose one-round cost is $E$, whereas its exact two-round and unrestricted deterministic costs are \[ \log_2 E+\tfrac12\log_2\log_2 E+O(1). \] Consequently, the largest nonadaptive-to-adaptive ratio is $(1+o(1))E/\log_2E$, including the optimal leading constant. A second restricted family exhibits a complete round hierarchy: its optimal $R$-round cost is $Θ(R E^{1/R})$. Thus the maximal gap already appears in two rounds, while other instances interpolate smoothly across all round budgets. Both constructions admit one-predicate realizations over the fixed alphabet $\{0,1,\#\}$ with logarithmic-length semantic spans and $O(E\log^2 E)$ total representation size. Under pointwise error $δ<1/2$ and worst-case expected cost, randomized nonadaptive complexity is exactly $(1-2δ)E$ for every instance with $E$ essential keys. Finally, for a fixed word $w$ and $h$ predicate names, the exact randomized minimax value is $(1-2δ)h sd(w)$, where $sd(w)$ counts distinct substring values; a span bound $s$ replaces $sd(w)$ by $sd_s(w)$. These results separate semantic information acquisition, parallel latency, and local symbolic matching cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑