发表机构
The University of Western Australia; Harvard University(西澳大利亚大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对RNA二级结构预测中矩阵到结构的转换问题,比较了四种提取算法,并提出可微分的SDSM归一化方法,可直接输出概率矩阵,在实验中优于基线,成为传统提取的可行替代。
AI 中文摘要
近年来,许多深度学习方法被提出用于RNA二级结构预测。这些方法通常输出一个权重矩阵 $W$,其中 $W_{ij}$ 是碱基 $i$ 与碱基 $j$ 配对时的任意权重。将该矩阵转换为预测的二级结构或碱基配对概率矩阵,通常涉及临时且有问题的不当下游算法。尽管这一转换步骤(我们称之为结构提取)非常重要,但在文献中受到的关注相对较少。在这项工作中,我们分析了训练方法与提取方法之间的一致性如何影响预测性能。为此,我们比较了四种提取算法:一种类似Nussinov的动态规划方法、最大权重图匹配以及SPOT-RNA和RiNALMo使用的贪心提取算法。这些算法在预训练的RiNALMo模型和本文训练的三种玩具模型上进行了评估:一种可微分的类似Nussinov模型、一种二元交叉熵(BCE)基线模型,以及一种在训练过程中结合了新颖的对称双随机矩阵(SDSM)归一化算法的模型,该模型能够直接输出碱基配对概率矩阵,而无需单独的提取步骤。这种SDSM归一化算法是可微分的,可以在训练和评估期间内联添加到任何深度学习模型中。我们发现,每种提取方法的性能在很大程度上取决于相应模型的训练方式。就玩具模型本身而言,SDSM模型表现出最强的整体性能:在所有四种提取算法下,它都优于BCE基线,并且产生的预提取输出最接近真实值。这些结果表明,SDSM归一化是传统结构提取的一种可行替代方案。
英文摘要
Many deep learning approaches to RNA secondary structure prediction have recently been proposed. They typically output a weight matrix $W$ where $W_{ij}$ is an arbitrary weight for base $i$ pairing with base $j$. Converting this matrix to a predicted secondary structure or base-pairing probability matrix typically involves ad hoc and problematic downstream algorithms. Despite the importance of this conversion step, which we refer to as structure extraction, it has received relatively little attention in the literature. In this work, we analyze how the congruence between training and extraction methods affects prediction performance. To do this, we compare four extraction algorithms: a Nussinov-like dynamic programming method, maximum-weight graph matching and the greedy extraction algorithms used by SPOT-RNA and RiNALMo. These are evaluated on outputs from the pretrained RiNALMo model and three toy models trained in this paper: a differentiable Nussinov-like model, a binary cross-entropy (BCE) baseline, and a model that incorporates a novel symmetric doubly stochastic matrix (SDSM) normalization algorithm during training which allows it to output base-pairing probability matrices directly, without a separate extraction step. This SDSM normalization algorithm is differentiable and can be added inline to any deep learning model during training and evaluation. We find that the performance of each extraction method depends strongly on how the corresponding model was trained. Considering the toy models themselves, the SDSM model showed the strongest overall performance: it outperformed the BCE baseline under all four extraction algorithms and produced pre-extraction outputs closest to the ground truth. These results suggest that SDSM normalization is a tractable alternative to traditional structure extraction.