AI 中文总结
该研究针对双块变量数据,提出三因子分解的非负降秩回归方法,实现响应与协变量变量组的共聚类,并通过四个数据集验证了其有效性。
AI 中文摘要
双块数据(即对同一观测对象测量得到的两组变量,例如微生物类群与代谢物)会引出协变量变量组如何与响应变量组相关联的问题。共聚类可针对单个矩阵解决该问题,但不适用于两个变量块;现有的双块方法要么仅对一侧进行聚类,要么返回带符号因子而非聚类结果。从多元线性回归 $Y_1\approx M Y_2$ 出发,我们对其非负系数矩阵进行三因子分解 $M=X_1\Theta X_2$(即三因子非负矩阵分解,tri-NMF),使得 $X_1$ 对响应变量进行软聚类,$X_2$ 对协变量进行软聚类,而 $\Theta$ 为经过检验的块对应矩阵。这使得该方法成为降秩回归(RRR)家族的非负成员,将 RRR 的低秩类别以基于部分的基的形式表达,如同非负矩阵分解(NMF)与主成分分析(PCA)的关系;该约束仅会限制拟合效果,因此预测精度并非目标,共聚类与经过检验的对应关系才是核心目标。我们给出乘法更新规则,通过交叉验证选择两个秩,并在基选择后为 $\Theta$ 开发条件 Wald 检验;当基固定时,检验大小为名义值,重新估计后为保守值,在响应变量相关时略高于名义值,而非零路径的幅度则在估计基的条件下保持稳定。我们在四个数据集上验证了该方法——即三因子分解的非负 RRR(NMF-RRR),涵盖置换结构(Doubs 群落生态学数据)、$p>n$ 下的弱交叉结构(nutrimouse 营养基因组学数据)、经筛选的微生物组-代谢组研究中的显著结构(FRANZOSA 数据,其中两个微生物类群与每个代谢物模块存在联合关联),以及分类特例(Wine 数据)。
英文摘要
Two-block data---two sets of variables measured on the same individuals, such as microbial taxa and metabolites---raise the question of how \emph{groups} of covariate variables relate to \emph{groups} of response variables. Co-clustering answers this for a single matrix, not for two variable blocks; existing two-block methods either cluster only one side or return signed factors rather than clusters. Starting from the multivariate linear regression $Y_1\approx M Y_2$, we give its non-negative coefficient matrix a tri-factorization $M=X_1ΘX_2$ (a tri-NMF), so that $X_1$ softly clusters the response variables, $X_2$ the covariate variables, and $Θ$ is a tested matrix of block correspondences. This makes the method the non-negative member of the reduced-rank regression (RRR) family, expressing RRR's low-rank class in a parts-based basis as NMF relates to PCA; the constraint can only restrict the fit, so predictive accuracy is not the aim; the co-clustering and tested correspondences are. We give multiplicative update rules, choose the two ranks by cross-validation, and develop a conditional Wald test for $Θ$ applied after basis selection; its size is nominal with fixed bases, conservative after re-estimation, and slightly above nominal under correlated responses, while a non-zero path's \emph{magnitude} stays conditional on the estimated bases. We illustrate the method---a tri-factorized non-negative RRR (NMF-RRR)---on four data sets spanning a permutation structure (Doubs, community ecology), a weak cross-structure under $p>n$ (nutrimouse, nutrigenomics), a pronounced one in a screened microbiome--metabolome study (FRANZOSA, where two microbial groups are jointly associated with each metabolite module), and a classification special case (Wine).