arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

何时准确率才是证据?泛化、验证与信息融合的统一理论

When Is Accuracy Evidence? A Unified Theory of Generalisation, Validation, and Information Fusion

JM Gorriz

arXiv 2610.03465首次发表:更新:

发表机构

Data Science and Computational Intelligence Institute; University of Granada; University of Cambridge(数据科学与计算智能研究所; 格拉纳达大学; 剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出gamma-CUBV统一框架,将K折交叉验证的准确率转化为真实风险上界,分离性能、不确定性与依赖性,为异质小样本提供原则性验证标准。

AI 中文摘要

K折交叉验证(CV)被广泛用作样本外性能的证据,尽管在异质数据下,各折既不是独立实验,也不是同等信息量的。交叉上界验证(CUBV)用真实风险上的保守上界取代了逐点CV准确率。在此,我们通过一个单一的指数框架推广CUBV,在该框架中,泛化差距的矩生成函数由累积量包络gamma(lambda)控制。这产生了一系列风险界,涵盖Hoeffding、Bernstein、依赖感知、PAC-Bayesian以及异质源融合等情形。对于K折CV,折间差距的依赖性通过联合次高斯代理矩阵建模。在等相关条件下,这给出了有效折数Keff = K/[1+(K-1)rho],表明当折间强相关时,增加K并不必然增加统计证据。该框架还扩展到预测器上的后验分布和加权多源融合,其中权重通过最小化未来风险的上界而非仅凭经验误差来选择。在异质多模态高斯混合上的训练线性分类器实验比较了K折CV与全样本重替换加风险校正。通过覆盖率和紧度评估各界的表现。在低维小样本设置中,K折划分可能增加不确定性,因为个别折对少数模态代表性不足,而校正后的重替换可以保持有效且更紧;这种效应随着样本量增加而消失。总体而言,gamma-CUBV将观测性能、不确定性、依赖性、模型复杂性和置信度分离为显式项,提供了一条从CV分数到风险陈述的统一路径,以及一个适用于神经影像等异质小样本应用的原则性验证标准。

英文摘要

K-fold cross-validation (CV) is widely used as evidence of out-of-sample performance, although folds are neither independent experiments nor equally informative under heterogeneous data. Cross Upper-Bound Validation (CUBV) replaces point-wise CV accuracy by conservative upper bounds on true risk. Here we generalise CUBV through a single exponential framework in which the moment-generating function of the generalisation gap is controlled by a cumulant envelope gamma(lambda). This yields a family of risk bounds covering Hoeffding-, Bernstein-, dependency-aware, PAC-Bayesian, and heterogeneous source-fusion settings. For K-fold CV, dependence between fold-wise gaps is modelled through a joint sub-Gaussian proxy matrix. Under equicorrelation, this gives an effective number of folds, Keff = K/[1+(K-1)rho], showing that increasing K does not necessarily increase statistical evidence when folds are strongly dependent. The framework is also extended to posterior distributions over predictors and weighted multi-source fusion, where weights are selected by minimising an upper bound on future risk rather than empirical error alone. Experiments with trained linear classifiers on heterogeneous multimodal Gaussian mixtures compare K-fold CV with full-sample resubstitution plus risk correction. Bounds are evaluated by coverage and tightness. In low-dimensional small-sample settings, K-fold partitioning can increase uncertainty because individual folds under-represent minority modes, while corrected resubstitution can remain valid and tighter; this effect disappears as sample size increases. Overall, gamma-CUBV separates observed performance, uncertainty, dependence, model complexity, and confidence into explicit terms, providing a unified route from CV scores to risk statements and a principled validation criterion for heterogeneous small-sample applications such as neuroimaging.

Comments52 pages, 30 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑