AI 中文总结
本文针对线性回归中假设模型与真实模型偏差导致的样本精度差异问题,提出了包含自变量及其函数的近似级数选择方法,以确定代表性样本量。
AI 中文摘要
通常,用于构建回归模型的样本所得到的分析与预测精度(包括复相关系数和残差均值),与另一同质样本的精度并不相等,实际上,基于另一样本的分析与预测精度会差得多。这一现象可由假设模型与真实模型之间的偏差来解释,该偏差由近似级数项数过多导致,近似级数用样本噪声描述所研究的现象。为过滤该噪声,需正确选择近似级数的项并确定其数量,而项数又取决于样本量。该近似级数不仅可包含自变量,还可包含自变量值的各种函数(如平方、立方、算法等)。本文给出了近似级数的选择方法。
英文摘要
Very often, accuracy of analysis and forecasting (multiple coefficient of regression and residual means) obtained for a sample used to formulate a regression model is not equal to the accuracy achieved for another homogeneous sample. Indeed, accuracy of analysis and forecasting based on another sample is much worse. This is explained by the discrepancy between a postulated and a real model. This discrepancy is caused by the redundancy of the number of terms of an approximating series. It describes a phenomenon under study with the sample noise. To filter this noise it is necessary to correctly choose the terms of an approximating series and to determine their number, which in turn depends on sample size. It is possible to include in this approximating series not only independent variables but also various functions of the value of the independent variables (e.g., squares, cubes, algorithms, etc.). This paper gives the procedure for selection of an approximating series.