arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16553cs.LG

蛋白质接触图上的离散Ricci曲率用于轻量级折叠分类

Discrete Ricci Curvature on Protein Contact Graphs for Lightweight Fold Classification

Jianru Shen

首次发表
浏览论文内容

中文总结 AI 辅助

研究蛋白质折叠分类,将α碳接触图上的离散Ricci曲率作为轻量级结构描述符,与多种基线比较。结果显示其性能优于平均池化的ESM-2嵌入,与持久同调结合性能最强,为预训练蛋白质语言模型嵌入提供实用替代方案。

中文摘要 AI 辅助

蛋白质折叠分类可通过基于序列的表示或结构描述符来实现,但轻量级手工描述符与预训练蛋白质语言模型嵌入之间的直接比较仍然有限。我们研究了α碳接触图上的离散Ricci曲率作为折叠分类的轻量级结构描述符。每个蛋白质结构域由一个22维固定长度特征表示,该特征来自Ollivier-Ricci和Forman-Ricci边缘曲率分布的汇总统计和分位数。我们在CATH前10拓扑分类和ASTRAL 40%同一性SCOPe前10折叠基准上进行评估,与几何、接触图统计、持久同调以及平均池化的ESM-2(150M)基线进行比较。在两个数据集上,轻量级结构描述符均显著优于平均池化的ESM-2嵌入,在ASTRAL 40% SCOPe基准上性能差距更大。仅Ricci使用22维,即ESM-2基线维度的3.4%,且在两个数据集上已优于平均池化的ESM-2。将Ricci与持久同调相结合产生最强性能,在CATH上达到0.71的宏F1,在SCOPe上达到0.68,特征向量为112维。这些结果表明,在某些情况下,轻量级可解释图描述符为预训练蛋白质语言模型嵌入提供了一种实用的替代方案。

英文摘要

Protein fold classification can be approached via sequence-based representations or structural descriptors, but direct comparisons between lightweight handcrafted descriptors and pretrained protein language model embeddings remain limited. We investigate discrete Ricci curvature on Calpha contact graphs as a lightweight structural descriptor for fold classification. Each protein domain is represented by a 22-dimensional fixed-length feature derived from summary statistics and quantiles of Ollivier-Ricci and Forman-Ricci edge curvature distributions. We evaluate on CATH top-10 Topology classification and on the ASTRAL 40%-identity SCOPe top-10 Fold benchmark, comparing against geometry, contact-graph statistics, persistent homology, and mean-pooled ESM-2 (150M) baselines. On both datasets, lightweight structural descriptors substantially outperform mean-pooled ESM-2 embeddings, with a larger performance gap on the ASTRAL 40% SCOPe benchmark. Ricci alone uses 22 dimensions, or 3.4% of the ESM-2 baseline dimensionality, and already outperforms mean-pooled ESM-2 on both datasets. Combining Ricci with persistent homology yields the strongest performance, achieving macro-F1 of 0.71 on CATH and 0.68 on SCOPe with a 112-dimensional feature vector. These results identify a regime where lightweight interpretable graph descriptors offer a practical alternative to pretrained protein language model embeddings.

补充信息

↑