全四分体一致性测试:自适应重构、随机验证与常查询可测试性
Testing Full Quartet Consistency: Adaptive Reconstruction, Random Verification, and Constant-Query Testability
- National Taiwan Ocean University(国立台湾海洋大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对n个分类单元的全四分体拓扑系统,提出自适应与非自适应单侧错误测试器,改进了查询复杂度,还证明了相关下界,为系统发育树相关的属性测试提供了新的高效方案。
AI中文摘要:
我们研究了关于n个分类单元上已解析四分体拓扑全系统的稠密属性测试问题:判定该系统是否由系统发育树诱导,或是与所有树诱导系统的距离至少为ε。我们的主要结果是一个显式多项式时间自适应单侧错误测试器,它通过锚定四分体查询重构候选树,并利用均匀随机四分体查询验证该候选树。在错误概率为δ时,该测试器使用O(n log n + ε⁻¹ log(1/δ))次查询。我们还给出了一种非自适应缓存锚定变体,使用C(n-1,3) + O(ε⁻¹ log(1/δ))次查询。这两种方法均改进了此前显式O(n³/ε)的查询复杂度界。由于输入包含C(n,4)=Θ(n⁴)个四分体条目,对于固定的ε和δ,两种测试器均使用o(C(n,4))次查询。我们还将全四分体系统在重标记等变下编码为有向三色4元结构,由此通过遗传有向超图测试得到一个不依赖n的单侧错误测试器,尽管其对ε的依赖在数量级上不实用。最后,我们证明了下界:在普通属性测试中,当n→∞时,每个自适应随机测试器(即使是双侧错误)渐近至少需要ln((1-δ)/δ)/ln(1/(1-ε))次查询;每个单侧错误测试器需要ln(1/δ)/ln(1/(1-ε))次查询,与随机验证项仅差舍入。对于更强的重构或拒绝任务,我们的上界在常数因子内是最优的:自适应复杂度为Θ(n log n + ε⁻¹ log(1/δ)),非自适应复杂度为Θ(n³ + ε⁻¹ log(1/δ))。
英文摘要:
We study dense property testing for full systems of resolved quartet topologies on $n$ taxa: determining whether a system is induced by a phylogenetic tree or is $\varepsilon$-far from every tree-induced system. Our main result is an explicit polynomial-time adaptive one-sided-error tester. It reconstructs a candidate tree through anchored quartet queries and verifies the candidate using uniformly random quartet queries. With error probability $δ$, it uses $O\!\left(n\log n+\varepsilon^{-1}\log(1/δ)\right)$ queries. We also give a non-adaptive cached-anchor variant using $\binom{n-1}{3}+O\!\left(\varepsilon^{-1}\log(1/δ)\right)$ queries. Both improve the previous explicit $O(n^3/\varepsilon)$ query bound. Since the input contains $\binom{n}{4}=Θ(n^4)$ quartet entries, both testers use $o\!\left(\binom{n}{4}\right)$ queries for fixed~$\varepsilon$ and~$δ$. We additionally encode full quartet systems, equivariantly under relabeling, as directed, three-colored $4$-ary structures. Hereditary directed-hypergraph testing then yields an $n$-independent one-sided-error tester, although its dependence on $\varepsilon$ is quantitatively impractical. Finally, we prove lower bounds. In ordinary property testing, every adaptive randomized tester, even with two-sided error, requires asymptotically at least $\ln((1-δ)/δ)/\ln(1/(1-\varepsilon))$ queries as $n\to\infty$. Every one-sided-error tester requires $\ln(1/δ)/\ln(1/(1-\varepsilon))$ queries, matching the random-verification term up to rounding. For the stronger reconstruct-or-reject task, our upper bounds are optimal up to constant factors: the adaptive and non-adaptive complexities are $Θ(n\log n+\varepsilon^{-1}\log(1/δ))$ and $Θ(n^3+\varepsilon^{-1}\log(1/δ))$, respectively.