arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Hyper-Fold:通过超图建模探索蛋白质序列-几何学习的表达极限

Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling

Yifan Feng, Guanjie Cheng, Shihui Ying, Shaoyi Du, Yue Gao

arXiv 2608.29207首次发表:更新:

发表机构

Tsinghua University; Zhejiang University; Shanghai University; Xi’an Jiaotong University(清华大学; 浙江大学; 上海大学; 西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Hyper-Fold通过超图建模探索蛋白质序列-几何学习的表达极限,其相关变体在多项蛋白质结构相关任务中取得最优结果,且Hyper-Fold-Pocket性能优于对比模型,参数与延迟更优。

AI 中文摘要

蛋白质结构建模依赖单一计算基元:残基的序列内容(即残基是什么)与其三维几何位置(即残基在哪里)之间的交互。这类层的表达极限是什么?研究表明,内容-几何外积上的完全双线性算子(所有二阶交互的充分统计量)是表达上限,而主流几何图神经网络(GNN)的加性消息传递对内容-几何绑定具有不可知性。随后引入Hyper-Fold,这是一种秩为K的可分离卷积骨干网络,以消息传递的成本逼近该上限:每个半径邻域被组织为序列超边和接触超边,由边缘条件矩阵值算子调制,该算子被分解为K个带几何生成系数的学习基算子。在酶功能预测、折叠分类和配体结合位点检测任务中,Hyper-Fold及其分层变体Hyper-Fold-Deep在蛋白质特异性结构编码器中取得最优结果;Hyper-Fold-Pocket作为锚定集预测头,在不使用序列语言模型特征的情况下,在UniSite-DS和两个零样本基准上超越UniSite-3D,参数减少68倍,延迟降低4.8倍——这表明足够具表达力的3D骨干网络可恢复融合架构此前从进化规模预训练中获取的信息。

英文摘要

Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (three-dimensional geometry). What is the expressive limit of this layer class? We show that the complete bilinear operator over content-geometry outer products--the sufficient statistic of all second-order interactions--is the expressive ceiling, while the additive message passing of mainstream geometric GNNs is provably blind to content-geometry binding. We then introduce Hyper-Fold, a rank-K separable convolutional backbone approaching this ceiling at message-passing cost: each radius neighborhood is organized into a sequence hyperedge and a contact hyperedge, modulated by an edge-conditioned matrix-valued operator factorized into K learned basis operators with geometry-generated coefficients. Across enzyme function prediction, fold classification, and ligand binding site detection, Hyper-Fold and its hierarchical variant Hyper-Fold-Deep achieve the best results among protein-specific structure encoders; Hyper-Fold-Pocket, an anchored set-prediction head, surpasses UniSite-3D on UniSite-DS and two zero-shot benchmarks with no sequence language model features, 68x fewer parameters, and 4.8x lower latency--suggesting that a sufficiently expressive 3D backbone recovers information that fusion architectures previously borrowed from evolution-scale pretraining.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑