发表机构
A*STAR Institute of Advanced Intelligence and Computing; RIKEN Center for Computational Science; Institute of Computing Technology, Chinese Academy of Sciences; The Chinese University of Hong Kong(新加坡科技研究局先进智能与计算研究所; 理化学研究所计算科学中心; 中国科学院计算技术研究所; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SparkleDock框架,通过优化GSO并行性、适配Tensor Core及负载均衡调度,在GPU上大幅加速大分子对接,实现512 GPU下数秒完成大规模高保真虚拟筛选。
AI 中文摘要
柔性大分子对接可对生物分子相互作用进行高保真预测,但大规模应用时成本过高。现有方法中,LightDock利用萤火虫群优化(Glowworm Swarm Optimization,GSO)实现精度,却存在并行性有限、计算不规则、负载失衡严重的问题,无法在GPU超级计算机上高效运行。本文提出SparkleDock,一种基于GSO的可扩展对接框架,支持近实时柔性对接。我们重新设计GSO以在萤火虫智能体层面暴露大量细粒度并行性,并将主导的能量评分计算重构为张量核心(Tensor Core)兼容的形式,通过结构化矩阵运算高效执行不规则的成对相互作用。我们还引入性能模型驱动的调度策略,实现跨GPU的负载均衡与外核扩展。SparkleDock在单张A100和H100 GPU上分别实现比LightDock高9.7倍和18.9倍的加速比,大规模部署时加速超两个数量级;在512个GPU上,它将对接时间从数小时缩短至数秒,实现了柔性对接此前无法实现的大规模高保真虚拟筛选。
英文摘要
Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among existing approaches, LightDock leverages Glowworm Swarm Optimization (GSO) for accuracy, yet suffers from limited parallelism, irregular computation, and severe load imbalance, preventing efficient execution on GPU supercomputers. We present SparkleDock, a scalable GSO-based docking framework enabling near-real-time flexible docking. We redesign GSO to expose massive fine-grained parallelism at the glowworm-agent level, and restructure the dominant energy scoring computation into a Tensor Core-compatible formulation, enabling efficient execution of irregular pairwise interactions through structured matrix operations. We further introduce a performance-model-driven scheduling for load balancing and out-of-core scaling across GPUs. SparkleDock achieves 9.7 $\times$ and 18.9 $\times$ speedups over LightDock on single A100 and H100 GPU, and delivers over two orders of magnitude acceleration at scale. On 512 GPUs, it reduces docking time from hours to seconds, enabling large-scale, high-fidelity virtual screening previously impractical with flexible docking.
CommentsAccepted in the International Conference for High Performance Computing, Networking, Storage, and Analysis(SC'26)