针对所有有序数据的扩展秩回归
Extended rank regression for all ordinal data
浏览论文内容
中文总结 AI 辅助
研究针对所有有序数据的回归模型,采用基于扩展秩概念的伪似然方法,能适应多种均值 - 方差关系和数据类型,无需相关估计与指定,该方法在极端数据情况无渐近信息损失,还可获贝叶斯参数估计等,共形校准能保证区间覆盖。
中文摘要 AI 辅助
回归模型的推断准确性很大程度上取决于模型对结果均值与方差关系的表示程度。由于这种关系很少是直接关注的对象,将其视为一个干扰参数而非尝试估计是很自然的。我们在单调变换线性回归模型的背景下采用这种方法,使用基于扩展秩概念的伪似然。此方法能适应广泛的均值 - 方差关系和任何有序数据类型,包括连续和离散有序数据,无需估计或预先指定变换,也无需决定将结果视为连续或离散。我们表明,扩展秩似然在连续和二元数据的两个极端情况下不会产生渐近信息损失,并且基于秩的预测区间可以在给定特征的情况下获得近似覆盖控制。通过简单的吉布斯采样算法可获得贝叶斯参数估计和预测区间。对于模型存疑的设置,贝叶斯预测分布的共形校准提供具有保证边际频率覆盖的区间。
英文摘要
The accuracy of inference from a regression model depends largely on how well the model represents the relationship between the mean and variance of the outcomes. As this relationship is rarely of direct interest, it is natural to treat it as a nuisance parameter, rather than attempt to estimate it. We take this approach in the context of a monotonically transformed linear regression model using a pseudo-likelihood based on an extended notion of ranks. This approach can accommodate a wide range of mean-variance relationships and any ordinal data type, including continuous and discrete ordered data, and requires no estimation or prior specification of the transformation, or decision to treat an outcome as continuous or discrete. We show that the extended rank likelihood incurs no asymptotic information loss at the two extremes of continuous and binary data, and that rank-based prediction intervals can obtain approximate coverage control conditional on the features. Bayesian parameter estimates and prediction intervals are available via a simple Gibbs sampling algorithm. For settings where the model is in doubt, conformal calibration of the Bayesian predictive distribution provides intervals with guaranteed marginal frequentist coverage.