arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高斯过程设计测试能做什么

What Can a Gaussian Process Design Test

Ivan De Boi, Marnix Van Soom

arXiv 2610.10122首次发表:更新:

发表机构

University of Antwerp; Vrije Universiteit Brussel(安特卫普大学; 布鲁塞尔自由大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出从设计阶段前瞻性解读高斯过程模型测试:通过核矩阵零空间和Gale对偶性区分结构性证据与先验证据,并用预测功效指导输入选择,提升对局部差异的检测能力。

AI 中文摘要

高斯过程(GP)模型与数据一致可能有两个原因:其假设是正确的,或者所选的输入永远无法表明这些假设是错误的。这种区别可以在观察到任何响应之前从设计中进行检查。每个模型都隐含了其无噪声响应在所选输入处必须满足的关系,例如中间值位于通过其两个相邻点的直线上。对于由有限特征构建的高斯过程,这些关系恰好是核矩阵的零空间。Gale对偶性赋予它们几何解释,其中每个观测对应一个向量,而能暴露错误的最小观测组是回路。对于其他核,这些关系变得软性:响应模式在先验下可能是不太可能的,而非代数上不可能。标准测试随后结合两类证据。结构性证据来自被违反的关系,并随着噪声下降而无限制增长。基于先验的证据仅表明偏离在先验下是不太可能的。例如,当所有输入位于区间两端时,高斯过程可以拒绝直线而支持大曲率,但这只是因为隐含的截距不太可能,而绝不是因为观察到了曲率。在模拟中,预测功效与观测拒绝率相匹配。通过预测功效选择下一个输入,将针对局部差异的功效从0.48提高到0.72,而按预测方差选择时为0.51,并且二维网格包含了拉丁超立方体所缺乏的可加性精确测试。该测试本身是经典的。贡献在于对该测试的前瞻性解读:在观察响应之前,设计已经决定了它能产生何种矛盾。

英文摘要

A Gaussian process (GP) model can agree with the data for two reasons: its assumptions are right, or the chosen inputs could never have shown that they are wrong. The distinction can be checked from the design before any responses are observed. Every model implies relations that its noiseless responses must satisfy at the chosen inputs, such as the middle value lies on the line through its two neighbours. For GPs built from finitely many features, these relations are exactly the null space of the kernel matrix. Gale duality gives them a geometric interpretation, in which each observation has a vector and the smallest groups of observations that can expose an error are the circuits. For other kernels the relations become soft: response patterns may be improbable under the prior rather than algebraically impossible. A standard test then combines two kinds of evidence. Structural evidence comes from a violated relation and grows without limit as the noise falls. Prior-based evidence only says that a departure is improbable under the prior. With all inputs at the two ends of an interval, for example, a GP can reject a straight line against a large curvature, but only because the implied intercept is improbable, never because curvature was seen. In simulations the predicted power matched the observed rejection rates. Choosing the next input by predicted power raised the power against a localised discrepancy from 0.48 to 0.72, against 0.51 when choosing by predictive variance, and a grid in two dimensions contained exact tests of additivity that a Latin hypercube lacked. The test itself is classical. The contribution is the prospective reading of that test: before observing the responses, the design already determines what kind of contradiction it can produce.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑