FoldKit:用于共折叠预测高效存储与检索的Python库
FoldKit: A Python library for efficient storage and retrieval of co-folding predictions
浏览论文内容
中文总结 AI 辅助
FoldKit是一款Python库,可高效存储检索AlphaFold 3共折叠预测结果,能降低5-15倍存储需求,支持多类生物分子相互作用数据集,助力大规模生物分子相互作用计算研究。
中文摘要 AI 辅助
AlphaFold 3(AF3)可通过对多个相互作用分子进行共折叠来实现生物分子复合物的结构预测,使其在从头蛋白质设计以及蛋白质-蛋白质、蛋白质-肽和其他生物分子相互作用的大规模研究中愈发有用。然而,系统性共折叠实验会产生大量输出数据,尤其是当每个输入复合物生成多个随机种子和样本时。我们推出FoldKit,这是一款用于高效存储和分析大规模AF3共折叠结果的Python包。FoldKit将原始AF3输出转换为紧凑的结构化表示,同时保留下游分析所需的元数据。FoldKit Python库提供对全局、单链及界面置信度指标的便捷编程访问,这些指标包括pLDDT、pTM、ipTM、ipAE和ipSAE,还提供集合级界面,用于访问和聚合单个输入在多个种子和样本上的这些指标。我们在三类AF3共折叠数据集上对FoldKit进行了基准测试:(i)每个输入含2条链的蛋白质设计数据集;(ii)每个输入含4条链的TCR-pMHC数据集;(iii)每个输入最多含22条链的pooled-AF3蛋白质-蛋白质相互作用数据集。我们发现,FoldKit根据数据集组成,与原始AF3输出相比可将存储需求降低约5-15倍,同时保持对单个预测、集合及置信度指标的直接编程访问。通过降低存储需求并促进对相关输出的编程访问,FoldKit推动了生物分子相互作用的大规模计算研究。FoldKit可从PyPI获取,可通过pip安装。
英文摘要
AlphaFold 3 (AF3) enables structure prediction of biomolecular complexes through co-folding multiple interacting molecules, making it increasingly useful for de novo protein design and for large-scale studies of protein-protein, protein-peptide, and other biomolecular interactions. However, systematic co-folding experiments can produce large volumes of output data, particularly when multiple random seeds and samples are generated for each input complex. We introduce FoldKit, a Python package for efficient storage and analysis of large-scale AF3 co-folding results. FoldKit converts raw AF3 outputs into a compact, structured representation while preserving the metadata needed for downstream analysis. The FoldKit Python library provides convenient programmatic access to global, single chain, and interface confidence metrics such as pLDDT, pTM, ipTM, ipAE, and ipSAE, as well as an ensemble-level interface for accessing and aggregating these metrics for a single input across multiple seeds and samples. We benchmark FoldKit on three types of AF3 co-folding datasets: (i) a protein design campaign with 2 chains per input, (ii) a TCR-pMHC dataset with 4 chains per input, and (iii) a pooled-AF3 protein-protein interaction dataset with up to 22 chains per input. We find that FoldKit reduces storage requirements by approximately 5-15-fold compared to native AF3 outputs, depending on dataset composition, while maintaining direct programmatic access to individual predictions, ensembles, and confidence metrics. By reducing storage requirements and facilitating programmatic access to relevant outputs, FoldKit facilitates large-scale computational studies of biomolecular interactions. FoldKit is available from PyPI and can be installed using pip.
发表机构
- Weill Cornell Medicine(威尔康奈尔医学院)
- Rockefeller University(洛克菲勒大学)
- Memorial Sloan Kettering Cancer Center(纪念斯隆-凯特琳癌症中心)
- Sloan Kettering Institute(斯隆研究所)
机构由 AI 辅助整理,请以论文原文为准。