arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11917cs.LG

面向可扩展多输出高斯过程回归的因子图方法

A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression

发表机构埃因霍温理工大学 · 拉齐动力公司
查看机构详情
  • Eindhoven University of Technology(埃因霍温理工大学)
  • Lazy Dynamics B.V.(拉齐动力公司)

机构由 AI 辅助整理,请以论文原文为准。

Wouter W. L. Nuijten, Esther G. van Pelt, Albert Podusenko, İsmail Şenöz, Wouter M. Kouw

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出一种因子图方法实现可扩展多输出高斯过程回归,通过链上高斯消息传递降低计算成本,在低输入维度任务中性能接近精确方法,在电力时间序列预测中兼具准确率与线性扩展性。

中文摘要 AI 辅助

多输出高斯过程回归的计算复杂度随观测数与输出数的乘积呈立方增长,且当不同输入对应不同输出观测时,稠密核矩阵方法需要专门处理。我们将多输出高斯过程回归表示为Forney风格的因子图,其中最近邻链将固定的C个候选输入排序为一维序列。沿该链,隐式Matérn过程通过线性高斯转移因子演化,而核心gionalization的线性模型则通过确定性混合因子和各输出的标量观测因子,将L个隐式过程混合为D个输出。后验计算可简化为链上的精确高斯消息传递,在链构建后的计算成本为O(C(DL² + L³)),缺失观测会省略其局部因子而无需进行任何协方差矩阵重构。因此,该方法的扩展性取决于数据样本数量和缺失观测率,最适合低输入维度的候选集。我们在合成输入维度扫描任务和电力时间序列预测任务中,将该因子图方法与精确核矩阵基线、稀疏变分诱导点基线、最近邻基线进行对比。在低输入维度下,因子图后验与精确核矩阵后验的匹配度很高;随着输入维度增加,二者差距逐渐扩大,但仍与两种近似基线具有竞争力。在电力时间序列任务中,我们的因子图方法在预测准确率上与三种基线相当,且随数据点数量呈线性扩展,而精确核矩阵方法变得不可行,诱导点基线则明显更慢。

英文摘要

Multi-output Gaussian process regression scales cubically in the number of observations times outputs, and dense kernel-matrix methods need bespoke handling whenever different outputs are observed at different inputs. We express multi-output Gaussian process regression as a Forney-style factor graph in which a nearest-neighbor chain orders a fixed candidate set of $C$ inputs into a one-dimensional sequence. Along this chain, latent Matérn processes evolve through linear-Gaussian transition factors, while the linear model of coregionalization mixes $L$ latent processes into $D$ outputs through a deterministic mixing factor and per-output scalar observation factors. Posterior computation reduces to exact Gaussian message passing on the chain at cost $\mathcal{O}(C(DL^2 + L^3))$ after chain construction, and missing observations omit their local factor without any covariance-matrix restructuring. The formulation therefore scales in the number of data samples and in the rate of missing observations, while remaining best suited to candidate sets in low input dimension. We compare the factor-graph formulation against an exact kernel-matrix baseline, a sparse-variational inducing-point baseline, and a nearest-neighbor baseline on a synthetic input-dimension sweep and on electricity time series forecasting. At low input dimension the factor-graph posterior tracks the exact kernel-matrix posterior closely, and the gap grows gradually as input dimension increases while staying competitive with both approximate baselines. On the electricity time series our factor-graph formulation matches all three baselines in forecast accuracy while scaling linearly in the number of data points, where the exact kernel-matrix method becomes infeasible and the inducing-point baseline remains substantially slower.

补充信息

↑