arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27431cs.LGcs.AI

SE(3)-MeanFlow:李群上的少步蛋白质主链生成

SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups

  • Purdue University(普渡大学)
  • Vanderbilt University(范德堡大学)
  • Freie Universität Berlin(柏林自由大学)

机构由 AI 辅助整理,请以论文原文为准。

Yikun Bai, Binghang Lu, Yikai Liu, Elaheh Akbari, Soheil Kolouri, Linxuan Wang, Ping He, Shuchan Wang, Ruqi Zhang, Guang Lin

AI总结:

SE(3)-MeanFlow将MeanFlow扩展到李群几何,通过新训练目标实现少步蛋白质主链生成,性能优于多倍采样的流匹配基线,适配高通量设计需求。

AI中文摘要:

蛋白质主链的生成建模有望从头设计具有指定结构和功能特性的蛋白质。现有扩散模型和流匹配模型能在SE(3)^N上生成高质量主链,但推理需通过数百次网络评估数值积分ODE,每次评估涉及李群指数映射,这是高通量设计任务的瓶颈。我们提出SE(3)-MeanFlow,这是将MeanFlow从欧氏空间扩展到蛋白质帧李群几何的少步生成框架。其原生工作于李代数so(3)和R^3,推导了旋转和平移的闭式平均速度恒等式,提供免模拟训练目标;还引入SE(3) alpha-Flow目标,消除旋转分支的雅可比-向量积,作为预热阶段,之后训练切换到小t稳定的MeanFlow损失,用于预训练剩余阶段及基于整流的后训练。在蛋白质主链生成任务中,SE(3)-MeanFlow的性能与使用多倍采样步数的流匹配基线相当或更优,且在少步场景中优势更显著,整流使其在相同计算预算下表现领先,仅在多样性上有适度损失。

英文摘要:

Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Existing diffusion and flow-matching models produce high-quality backbones on SE(3)^N, but inference requires numerically integrating an ODE over hundreds of network evaluations, each involving a Lie group exponential map - a bottleneck for high-throughput design campaigns. We introduce SE(3)-MeanFlow, a few-step generative framework that extends MeanFlow from Euclidean space to the Lie group geometry of protein frames. Working natively in the Lie algebra so(3) and in R^3, we derive closed-form average-velocity identities for rotations and translations, giving simulation-free training targets. We further introduce an SE(3) alpha-Flow objective that removes the Jacobian-vector product from the rotation branch and serves as a warm-up stage, after which training switches to a small-t stabilized MeanFlow loss that is used for the remainder of pretraining and for rectification-based post-training. In protein backbone generation, SE(3)-MeanFlow matches or exceeds flow-matching baselines that use several times more sampling steps, and its advantage widens in the few-step regime, where rectification lets it lead at every matched budget - at a modest cost in diversity.

↑