arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种估计机器学习模型样本量的统计方法

A Statistical Approach to Estimating Sample Size of Machine Learning Models

Dat Phan-Trong, Sunil Gupta, Svetha Venkatesh

arXiv 2609.09547首次发表:更新:

发表机构

Deakin Applied Artificial Intelligence Initiative, Deakin University(迪肯大学应用人工智能计划)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器学习模型样本量确定困难的问题,提出用局部线性表示近似非线性模型并评估局部统计功效来估计样本量的框架。

AI 中文摘要

机器学习(ML)预测模型的样本量确定具有挑战性,因为传统的功效分析通常需要预先指定预测变量与结果之间的关系以及效应结构。非线性ML模型学习复杂的预测表面,这些表面不允许直接进行解析功效计算。我们提出一个框架,该框架用局部线性表示来近似非线性ML模型,并通过评估这些局部区域上的统计功效来估计样本量需求。

英文摘要

Sample size determination for machine learning (ML) prediction models is challenging because conventional power analysis typically requires the predictor-outcome relationship and effect structure to be specified a priori. Nonlinear ML models learn complex prediction surfaces that do not admit straightforward analytical power calculations. We propose a framework that approximates nonlinear ML models with localized linear representations and estimates sample size requirements by evaluating statistical power across these local regions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑