arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双向预算下的导数高斯过程

Derivative Gaussian Processes on a Two-Direction Budget

Hyunseok Seung, Matthias Katzfuss

arXiv 2610.10428首次发表:更新:

发表机构

University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种预算为每梯度两个方向的导数高斯过程,在Vecchia近似下用至多2m个方向导数表示md个梯度坐标,成本更低且精度可与精确方法媲美,甚至优于仅函数GP。

AI 中文摘要

梯度观测有望提高高斯过程(GP)替代模型的准确性,但长期以来,纳入梯度的成本一直阻碍着这一前景的实现。我们提出了一种导数高斯过程,其预算仅为每个观测梯度两个方向。一个方向专注于每个梯度对目标预测的直接贡献,而另一个方向则通过其与条件函数值的相关性来聚合其间接贡献。在Vecchia近似中,每个预测以$d$维空间中的$m$个邻近输入为条件,该构造用至多$2m$个方向导数表示其$md$个梯度坐标,从而每个预测目标的密集分解成本为$\mathcal{O}(m^3)$。对于一般的条件集,我们限制了相对于使用完整梯度的后验近似误差,并刻画了误差较小或近似精确的条件。在模拟中,我们的方法在相同条件集大小下与领先的精确梯度缩减方法的准确性相匹配。由于其成本随条件集大小的增长要慢得多,它可以使用远超精确方法内存限制的条件集,以一小部分时间和内存达到更低的预测误差。值得注意的是,我们的方法可以利用梯度观测,同时所需的计算时间或内存少于仅函数的GP基线。

英文摘要

Gradient observations promise more accurate Gaussian process (GP) surrogates, but the cost of incorporating them has long stood in the way of realizing that promise. We propose a derivative GP with a budget of just two directions per observed gradient. One direction focuses on each gradient's direct contribution to target prediction, while the other aggregates its indirect contributions through correlations with the conditioning function values. Within a Vecchia approximation, where each prediction conditions on $m$ nearby inputs in $d$ dimensions, this construction represents their $md$ gradient coordinates using at most $2m$ directional derivatives, giving $\mathcal{O}(m^3)$ dense factorization cost per prediction target. For general conditioning sets, we bound the posterior approximation error relative to using full gradients and characterize when the error is small or the approximation is exact. In simulations, our method matches the accuracy of a leading exact gradient-reduction method at equal conditioning set size. Because its cost grows much more slowly with that size, it can use conditioning sets well beyond the memory limit of the exact method, reaching lower prediction error with a small fraction of the time and memory. Notably, our method can exploit gradient observations while requiring less computation time or memory than function-only GP baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑