高阶马尔可夫模型的大样本性质
Large Sample Properties of Higher Order Markov Models
AI总结:
本文研究阶数随序列长度增长的有限字母表高阶马尔可夫链大样本性质,通过嵌入一阶链与返回时间分解建立加性泛函的中心极限定理,还在二元VLMC中推导显式下界给出增长条件,为稀疏高阶模型的推断提供渐近基础。
AI中文摘要:
我们研究有限字母表$\Sigma$上阶数$m_n$可随序列长度$n$增长的高阶马尔可夫链的大样本性质。通过将该过程嵌入到$\Sigma^{m_n}$上的一阶链中,并利用返回时间分解,我们在自然的遍历性和稀疏性条件下,建立了加性泛函$\sum_{t}\\! g_n(Y_t^{(n)})$的中心极限定理(CLT)。归一化项涉及到适当选取状态的平稳返回时间,且适用于满足$m_n\\!\to\\!\infty$和$m_n/n\\!\to\\!0$的三角阵列。我们进一步在二元变长马尔可夫链(VLMC,variable length Markov chain)中阐释了这些假设,推导了平稳质量的显式下界,该下界给出了确保CLT成立的具体增长 regime(例如$m_n\log m_n/n \to 0$)。这些结果为稀疏/分块高阶模型(包括VLMC和有效维数随样本量增长的稀疏马尔可夫模型(SMM,sparse Markov models))中的推断提供了渐近基础。
英文摘要:
We study large-sample properties of higher-order Markov chains on a finite alphabet $Σ$ when the order $m_n$ is allowed to grow with the sequence length $n$. By embedding the process into a first-order chain on $Σ^{m_n}$ and exploiting return-time decompositions, we establish a central limit theorem for additive functionals $\sum_{t}\! g_n(Y_t^{(n)})$ under natural ergodicity and sparsity conditions. The normalization involves the stationary return time to a suitably chosen state and accommodates triangular arrays with $m_n\!\to\!\infty$ and $m_n/n\!\to\!0$. We further illustrate the assumptions in a binary variable length Markov chain (VLMC), deriving explicit lower bounds on stationary masses that yield a concrete growth regime (e.g., $m_n\log m_n/n \to 0$) ensuring the CLT. These results provide asymptotic foundations for inference in sparse/partitioned higher-order models; including VLMCs and sparse Markov models (SMMs) where the effective dimensionality grows with the sample size.