对代数而非张量化进行评分:动力系统转移算子模型的降维
Score the Algebra, Not the Span: Dimension Reduction for Transfer Operator Models of Dynamical Systems
浏览论文内容
中文总结 AI 辅助
该研究针对多弱相互作用组件的动力系统,提出对σ-代数而非张量化评分的降维方法,可避免线性掩码问题,用少量代数坐标恢复被遗漏组件,实现低样本预测。
中文摘要 AI 辅助
动力系统的降维是标准做法,标准途径是谱方法:通过主导模态对转移(Koopman)算子进行建模。我们证明,在由几个弱相互作用组件构成的系统中——这是物理和生物环境中的常见结构——这种方法要么需要指数级数量的模态,要么会丢失整个组件:该组件未出现在模型中,而非被粗略建模,且其任何函数都无法以任何精度被预测。我们将此现象称为线性掩码。其原因是基于秩的模型每个模态占用一个坐标。我们提议改为对坐标生成的σ-代数进行评分,这样乘积和幂次是免费的,组件的成本仅由其生成元决定,而非其所有相互作用。该准则是嵌入的当前状态与未来状态之间的χ²散度,它带有预算保证:动力学的固有维度的两倍坐标足以构建一个嵌入,其代数承载算子的全部谱及其完整的无限秩。在变分形式中,该准则可采用现成的估计器,将其评判器限制为双线性类会返回张量化上的VAMP评分,因此基于秩的方法是同一方法族的一端。我们在已发布基准系统的组合上演示了所提出的目标。我们展示了在所有秩k<100时,基于秩的方法完全遗漏了被掩码的组件,而十个代数坐标即可恢复所有组件。此外,所得的代数表示支持从少量标签预测被掩码的组件,而直接从高维观测或VAMP特征进行回归则失败。
英文摘要
Dimension reduction for dynamical systems is standard practice, and the standard route is spectral: model the transfer (Koopman) operator by its leading modes. We show that on systems assembled from several weakly interacting components --- a structure common in physical and biological settings --- this may either require an exponential number of modes, or drop an entire component: the component is absent from the model rather than modeled coarsely, and no function of it can be predicted at any accuracy. We call this linear masking. The cause is that a rank-based model pays one coordinate per mode. We propose to score instead the $σ$-algebra the coordinates generate, so that products and powers come free and a component's cost is governed only by its generators rather than by all its interactions. The criterion is a $χ^2$-divergence between the embedded present and future, and it carries a budget guarantee: twice the intrinsic dimension of the dynamics is enough coordinates for an embedding whose algebra carries the operator's entire spectrum, with its full infinite rank. In variational form the criterion admits off-the-shelf estimators, and restricting its critic to the bilinear class returns the VAMP score on the span, so rank-based methods are one end of the same family. We demonstrate the proposed objective on a composite of published benchmark systems. We exhibit examples where the rank-based methods completely miss the masked components at all ranks $k<100$, while ten algebra coordinates recover all of them. In addition, the resulting algebra representation supports predicting the masked components from few labels, while direct regression from the high-dimensional observation or from the VAMP features fail.
发表机构
- Technion, IIT(以色列理工学院)
- NVIDIA(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。