用于模式发现的非线性张量分解
Nonlinear Tensor Decomposition for Pattern Discovery
浏览论文内容
中文总结 AI 辅助
针对截断或饱和数据下线性CP分解不可靠的问题,提出非线性CP模型(NCP),采用ADMM框架求解,在合成和模拟代谢组学数据上更准确地恢复潜在因子。
中文摘要 AI 辅助
CANDECOMP/PARAFAC (CP) 分解被广泛用于从多路数据(也称为高阶张量)中揭示潜在模式。然而,当数据被截断或饱和时,线性 CP 模型变得不可靠。现有的几种补救措施包括将截断条目视为缺失或进行插补;然而,这两种方法在不同场景下均有不足。我们提出了一种非线性 CP 模型(NCP),用于在此设置下进行模式发现,将非线性矩阵分解的思想扩展到张量,并使用灵活的基于 ADMM(交替方向乘子法)的框架进行求解,该框架能够适应多种非线性。利用具有已知因子的合成数据和模拟代谢组学数据(其中截断源于仪器检测限),我们证明 NCP 比仅拟合观测条目或阈值插补数据的 CP 更准确地恢复了潜在因子。
英文摘要
The CANDECOMP/PARAFAC (CP) decomposition is widely used for revealing the underlying patterns from multiway data (also referred to as a higher-order tensor). When the data is clipped or saturated, however, the linear CP model becomes unreliable. Several remedies exist such as treating the clipped entries as missing or imputation; however, both fall short in different scenarios. We propose a nonlinear CP model (NCP) for pattern discovery in this setting, extending ideas from nonlinear matrix decompositions to tensors, and solve it using a flexible ADMM (Alternating Direction Method of Multipliers)-based framework that can accommodate a variety of nonlinearities. Using synthetic data with known factors and simulated metabolomics data where clipping arises from instrument detection limits, we demonstrate that NCP recovers the underlying factors more accurately than CP fitted to only the observed entries or threshold-imputed data.