发表机构
University of Warwick; ESSEC Business School(华威大学; 埃塞克商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究缓慢混合高斯隐马尔可夫模型中的簇恢复,揭示统计难度集中于罕见转移点,提出自适应多项式时间算法,达到尖锐恢复速率,有效复杂度由潜在转移数而非样本量决定。
AI 中文摘要
我们研究了在缓慢混合机制下的两状态高斯隐马尔可夫模型中的簇恢复问题,其中转移概率 $\delta$ 可能随样本量趋于零。时间上的持续性产生了长的同质段,并从根本上改变了相对于独立同分布高斯混合基准的恢复问题。标签依赖性并非纯粹不利,而是提供了可利用的结构:统计难度集中在罕见的转移点附近,从而产生了明确由其频率 $\delta$ 控制的恢复机制。我们首先刻画了离线与在线聚类的oracle贝叶斯风险,精确到通用常数。两个问题都表现出非标准的多项式风险机制,但序贯约束引入了额外的对数定位成本,产生了真正的在线/离线差距。我们还研究了带有延迟标签反馈的固定滞后预测,并确定了有效记忆范围,超过该范围后过去的标签不再改善最优速率。当信号方向未知且可能为高维时,问题由有效信号强度 \\[ r_n^2 = \frac{\\|\theta\\|^4/\sigma^4} {\\|\theta\\|^2/\sigma^2+d/n} \\] 控制。我们推导了极小极大下界,并构造了完全自适应的多项式时间过程,以达到尖锐的恢复速率。特别地,几乎完全恢复恰好当 $r_n^2\gg\delta$ 时可能。如果 $n\delta\gtrsim\log n$,对信号 $\theta$ 和转移概率 $\delta$ 的自适应不会产生一阶损失,且尖锐的精确恢复阈值(在汉明损失下)为 \\[ r_n^2 = 2\log(n\delta)\\,(1+o(1)). \\] 因此,问题的有效复杂度由潜在转移的数量决定,而非样本量本身。
英文摘要
We study cluster recovery in a two-state Gaussian hidden Markov model in the slowly mixing regime, where the transition probability $δ$ may vanish with the sample size. Temporal persistence creates long homogeneous segments and fundamentally changes the recovery problem relative to the i.i.d. Gaussian mixture benchmark. Rather than being purely adverse, label dependence provides exploitable structure: the statistical difficulty becomes concentrated near the rare transition points, giving rise to recovery regimes governed explicitly by their frequency $δ$. We first characterize, up to universal constants, the oracle Bayes risk for offline and online clustering. Both problems exhibit a nonstandard polynomial risk regime, but the sequential constraint induces an additional logarithmic localization cost, yielding a genuine online/offline gap. We also study fixed lag prediction with delayed label feedback and identify the effective memory horizon beyond which past labels no longer improve the optimal rate. When the signal direction is unknown and possibly high dimensional, the problem is governed by the effective signal strength \[ r_n^2 = \frac{\|θ\|^4/σ^4} {\|θ\|^2/σ^2+d/n}. \] We derive minimax lower bounds and construct fully adaptive polynomial time procedures attaining the sharp recovery rates. In particular, almost full recovery is possible precisely when $r_n^2\ggδ$. If $nδ\gtrsim\log n$, adaptation to both the signal $θ$ and the transition probability $δ$ incurs no first order loss, and the sharp exact recovery threshold (under the Hamming loss) is \[ r_n^2 = 2\log(nδ)\,(1+o(1)). \] Thus the effective complexity of the problem is governed by the number of latent transitions, rather than by the sample size itself.