AI 中文总结
针对长平稳时间序列,提出基于频域Whittle似然和分布式FFT的谱分治MCMC方法,实现可扩展贝叶斯推断,理论保证收敛性,实验优于时域分治方法。
AI 中文摘要
时间序列模型中固有的时间依赖性对分布式贝叶斯推断构成了根本性挑战,因为它排除了基于时间域朴素独立性假设的尴尬并行算法。我们提出了一种用于平稳时间序列可扩展贝叶斯推断的频域框架,该框架利用了Whittle似然背后的渐近独立性。为了利用并行计算资源,我们开发了一种分布式快速傅里叶变换,并将其与现代集群计算框架中的尴尬并行马尔可夫链蒙特卡罗(MCMC)算法集成。这使得能够分析超过单个计算节点内存容量或计算时间成为瓶颈的时间序列。所提出的方法通过与将分治算法应用于频域而非时域分区,兼容于一大类现有的针对独立数据的分治算法。我们证明了谱分治MCMC后验近似相对于精确时域后验的误差,在全数据Whittle后验模式的收缩邻域内依概率收敛到零。相应的收敛速度也被推导出来。在多个实验中,我们证明了我们的方法提供了对全数据Whittle后验的准确近似。所提出的方法被证明优于当前最先进的时域分治方法,特别是对于高度持久的过程。该方法进一步通过将半长程模型拟合到长时间气象时间序列来说明。
英文摘要
The temporal dependence inherent in time series models poses a fundamental challenge for distributed Bayesian inference, as it precludes embarrassingly parallel algorithms based on naive independence assumptions in the time domain. We propose a frequency-domain framework for scalable Bayesian inference in stationary time series that exploits the asymptotic independence underlying the Whittle likelihood. To exploit parallel computing resources, we develop a distributed fast Fourier transform and integrate it with embarrassingly parallel Markov chain Monte Carlo (MCMC) algorithms within a modern cluster-computing framework. This enables the analysis of time series that exceed the memory capacity of a single computational node or for which computation time is a bottleneck. The proposed methodology is compatible with a broad class of existing divide-and-conquer algorithms for independent data by applying them to frequency-domain rather than time-domain partitions. We establish that the error of the spectral divide-and-conquer MCMC posterior approximation relative to the exact time-domain posterior converges to zero in probability in a shrinking neighbourhood of the full-data Whittle posterior mode. The corresponding convergence rate is also derived. Across several experiments, we demonstrate that our approach provides accurate approximations to the full-data Whittle posterior. The proposed method is shown to outperform the current state-of-the-art time-domain divide-and-conquer methodology, particularly for highly persistent processes. The methodology is further illustrated by fitting a semi-long range model to a long meteorological time series.