arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03439cs.DScs.LGmath.COmath.PR

模式计数的质量控制算法

Quality Control Algorithms for Pattern Counting

Cassandra Marcussen, Ronitt Rubinfeld, Madhu Sudan

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对序列的模式计数与排列模式计数问题,利用质量控制问题定义的不对称性,给出了多项式时间的高效算法,同时证明了这类算法的查询复杂度下界。

中文摘要 AI 辅助

近期,Marcussen、Rubinfeld和Sudan引入了质量控制问题的概念,旨在判断给定输入是否真正随机。形式上,这类问题的目标是接受来自指定分布的典型输入,同时拒绝那些指定统计量值与分布基线偏差过大的输入,这对应了将指定统计量作为随机性质量代理的经验做法。然而,经验算法并未利用质量控制问题定义中的不对称性——这类问题需要最坏情况下的可靠性保证,而仅追求平均情况下的完备性。他们的工作抽象出了强调该不对称性的问题定义,并将其用于给出评估图随机性的高效质量控制算法。在本文中,我们引入并研究序列上的质量控制问题,其目标是区分独立同分布字符序列与某些指定模式作为子序列出现过于频繁(或过于稀少)的序列。我们在有限字母设置和实值序列两种场景下考虑该问题,前者称为模式计数问题,后者中模式的自然概念是子序列中字符的相对顺序,称为排列模式计数问题。对于长度为n的最坏情况序列,近似计数长度为k的(排列)模式的算法可证明需要对序列进行指数级于k的查询;相比之下,我们表明通过利用质量控制定义中的不对称性,可给出在poly(k)时间内运行的算法来解决这些问题。我们还证明,任何(在某些自然分布上的)质量控制算法需要超线性于k的查询。

英文摘要

In recent work, Marcussen, Rubinfeld, and Sudan introduced the notion of quality control problems, which aim to capture the task of determining if a given input is truly random. Formally, their goal is to accept typical inputs from the specified distribution while rejecting every input whose value of a specified statistic is far from the distributional baseline. This captures the empirical practice of using specified statistics as a proxy for the quality of randomness. Empirical algorithms, however, have not exploited the asymmetry in the definition of quality control problems, which require soundness guarantees in the worst-case while only seeking average-case completeness. Their work abstracted a problem definition emphasizing this asymmetry and used it to give efficient quality control algorithms for assessing the randomness of graphs. In this work, we introduce and study quality control problems over sequences, where the goal is to distinguish a sequence of i.i.d. characters from sequences where some specified pattern appears too often (or too infrequently) as a subsequence. We consider this problem in both the finite-alphabet setting and for real-valued sequences. We refer to the former setting as the pattern counting problem. In the latter case, the natural notion of a pattern is to consider the relative ordering of the characters in the subsequence, and we refer to this as the permutation pattern counting problem. Algorithms to approximately count (permutation) patterns of length $k$ in a worst-case sequence of length $n$ can provably require exponential in $k$ queries into the sequence. In contrast, we show that by taking advantage of the asymmetry in the definition of quality control, we give algorithms that run in poly$(k)$ time to solve these problems. We also prove that any quality control algorithm (over some natural distributions) requires superlinear queries in $k$.

发表机构

  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑