arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VARMA模型的可扩展估计

Scalable estimation of VARMA models

Daniel Paulin, Victor Elvira

arXiv 2608.06340首次发表:更新:

发表机构

Nanyang Technological University; University of Edinburgh(南洋理工大学; 爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种计算量与序列长度无关的VARMA模型估计框架,通过偏自相关重参数化等技术得到两种估计量,在多维度数据集上表现优于或匹配相关基线,解决了高维VARMA估计的计算难题。

AI 中文摘要

向量自回归移动平均(VARMA)模型长期以来被认为在中等维度以上不实用:其似然函数非凸,参数化仅在等价意义上可识别,且每次评估都需要遍历整个序列。然而,其移动平均项能用少量参数捕捉到纯自回归模型需多阶滞后才能匹配的内容。我们提出一种估计框架,消除了这一计算障碍:每次优化迭代的计算量与序列长度$T$无关。该框架结合了偏自相关重参数化(该重参数化从结构上保证平稳性与可逆性)、对重参数化系数的高斯先验(对角线与非对角线项采用不同尺度),以及仅通过固定大小充分统计量依赖数据的损失函数,该损失函数通过Parseval(傅里叶)恒等式在截断长度上以近线性成本计算。这产生了两个点估计量:正则化最小二乘拟合和协方差边缘化最大后验估计量。我们证明,在固定维度下,两者都能以近参数速率恢复真实过程的无限自回归表示,因此截断不会引入渐近偏差。相同机制以相同主导成本扩展到季节性动态、外生回归变量(VARMAX)和滚动窗口重新拟合。实验表明,从$d=10$到$d=40$时,这些估计量的预测误差接近先验估计的预测误差(而经典条件最大似然估计会返回不可逆拟合,其预测结果发散),并且在零售需求、气象和空气质量数据上,与VAR、贝叶斯VAR、逐分量ARMA和稀疏VARMA基线相比,表现相当或更优。这使得基于似然的VARMA估计在每次迭代成本与序列长度无关的情况下,适用于从业者迄今为止依赖VAR模型的问题规模。

英文摘要

Vector autoregressive moving-average (VARMA) models have long been considered impractical beyond moderate dimensions: the likelihood is non-convex, the parametrization is identified only up to equivalence, and every evaluation costs a pass over the entire series. Yet their moving-average term captures with a few parameters what a pure autoregression matches only with many lags. We introduce an estimation framework that removes this computational barrier: each optimization iteration is independent of the series length $T$. The framework combines a partial-autocorrelation reparametrization that guarantees stationarity and invertibility by construction, Gaussian priors on the reparametrized coefficients with separate scales for diagonal and off-diagonal entries, and losses that depend on the data only through fixed-size sufficient statistics, evaluated by a Parseval (Fourier) identity at near-linear cost in the truncation length. This yields two point estimators: a regularized least-squares fit and a covariance-marginalized maximum-a-posteriori estimator. We prove that both recover the infinite-autoregressive representation of the true process at a near-parametric rate in fixed dimension, so the truncation introduces no asymptotic bias. The same machinery extends, at the same leading cost, to seasonal dynamics, exogenous regressors (VARMAX), and rolling-window refits. Empirically, the estimators stay close to the oracle forecast error from $d=10$ to $d=40$ (where classical conditional MLE returns non-invertible fits whose forecasts diverge) and match or beat VAR, Bayesian-VAR, component-wise ARMA, and sparse-VARMA baselines on retail-demand, meteorological, and air-quality data. This brings likelihood-based VARMA estimation, at a per-iteration cost independent of the series length, to the problem sizes where practitioners have so far relied on VAR models.

Comments60 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑