发表机构
CMU; NUS(卡内基梅隆大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出无监督原子策略优化框架APO,通过双奖励机制实现原子系统三维结构预测,在晶体与抗体预测中优于全监督基线,提升了匹配率、结构保真度与推理效率。
AI 中文摘要
预测原子系统的三维结构是推动材料科学与药物发现的基础。尽管流匹配模型(如FlowDPO)近期在该领域展现出潜力,但其性能高度依赖通过监督偏好学习与真实坐标的对齐。然而,获取新晶相或从头设计蛋白质的实验标签成本极高,在数据稀缺场景下成为结构建模的瓶颈。本研究提出APO(Atomic Policy Optimization),一种完全无监督的对齐框架,无需真实参考结构。APO将组相对策略优化适配到三维原子环境,采用新颖的双奖励机制:(i)通过样本相似性的特征分解强化策略的主导潜在结构模式;(ii)确保热力学稳定性。该框架通过在采样组内识别物理合理构型,使模型实现“自校正”。在晶体与抗体结构预测的广泛基准测试中,APO始终优于全监督基线,在匹配率与结构保真度上达到新的SOTA;此外,APO可有效拉直概率路径,显著提升推理效率。结果表明,与有噪的监督坐标匹配相比,内在物理一致性可作为更优的对齐指导。
英文摘要
Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) have recently shown promise in this domain, their performance relies heavily on alignment with ground-truth coordinates via supervised preference learning. However, obtaining experimental labels for novel crystal phases or de novo proteins is prohibitively expensive, creating a bottleneck for structural modeling in data-scarce regimes. In this work, we propose (Atomic Policy Optimization), a fully unsupervised alignment framework that eliminates the need for ground-truth reference structures. APO adapts group-relative policy optimization to 3D atomic environments, utilizing a novel dual-reward mechanism: (i) a that reinforces the policy's dominant latent structural modes through eigen-decomposition of sample similarities, and (ii) a that enforces thermodynamic stability. Our framework enables the model to ``self-correct'' by identifying physically plausible configurations within sampled groups. Extensive benchmarks on crystal and antibody structure prediction demonstrate that APO consistently outperforms fully supervised baselines, achieving a new state-of-the-art in match rates and structural fidelity. Furthermore, we show that APO effectively straightens probability paths, significantly improving inference efficiency. Our results suggest that intrinsic physical consistency can serve as a superior guide for alignment compared to noisy, supervised coordinate matching.