AI 中文总结
研究非负矩阵分解中最优秩选择问题,通过在因子矩阵引入列$\ell_{2,0}$-范数正则化,采用热启动$\lambda$路径及开发多种方法,经收敛分析和实验验证,该框架能有效估计秩,兼具精度与效率。
AI 中文摘要
确定最优分解秩是非负矩阵分解中一个基本但极具挑战性的问题,传统上依赖启发式阈值或计算成本高昂的交叉验证。我们在两个因子矩阵上引入列$\ell_{2,0}$-范数正则化以促进秩降低。从高估的分解维度开始,采用热启动的$\lambda$路径逐步消除冗余分量并产生数据驱动的秩估计。对于由此产生的非凸和不连续优化问题,我们开发了惯性近端交替线性化最小化方法、其尺度平衡变体以及基于P平稳性的近端活动集方法。尺度平衡策略控制配对因子列之间的数值不平衡同时保留重构矩阵。我们建立了所提出模型恢复潜在非负秩的条件,并基于Kurdyka-Łojasiewicz进行收敛分析,表明在所陈述的算法条件下,所提出方法生成的整个序列收敛到临界点。在合成和基准数据集上的数值实验以及与奇异值硬阈值和交叉验证的比较表明,所提出的框架在测试设置中提供了良好的秩估计精度和计算效率。
英文摘要
Nonnegative matrix factorization represents nonnegative signals as additive combinations of latent components, but its factorization rank, and hence the model order, must usually be specified beforehand. An underestimated order discards signal structure, whereas an overestimated order produces redundant components and unstable decompositions. We propose a column $\ell_{2,0}$-regularized formulation that estimates the model order from an initial upper bound by suppressing inactive columns in both factors. A warm-started regularization path progressively removes redundant components without changing the factor dimensions, and a marginal reconstruction-loss criterion selects an order along the path. To solve the resulting nonconvex and discontinuous problem, we develop an inertial proximal alternating linearized minimization method, a scale-balanced variant, and a proximal active-set method based on P-stationarity. The balancing operation equalizes the norms of paired factor columns while preserving their rank-one products. We characterize the critical points and local minimizers of the model, provide a sufficient-condition result for rank recovery, and prove whole-sequence convergence of the proposed algorithms under explicit step-size and inertial-parameter conditions using the Kurdyka--Łojasiewicz framework. Dedicated experiments show that the warm-started $λ$-path is more efficient than increasing- and decreasing-order discrete $r$-paths, while scale balancing yields a more stable rank-selection path. Experiments on synthetic data and diverse signal benchmarks show that both iPALM and PASM provide reliable model-order estimates with favorable computational efficiency.
Comments17 pages, 4 figures