FAR-DPO:用于环肽设计的可行性感知与鲁棒直接偏好优化
FAR-DPO: Feasibility-Aware and Robust Direct Preference Optimization for Cyclic Peptide Design
AI总结:
FAR-DPO是一种架构无关的环肽设计框架,通过可行性感知偏好构建与难度感知分组鲁棒优化,在CPSea LNR基准测试中提升了环肽设计的可行性与结合性能。
AI中文摘要:
环肽因具有高结合亲和力和结构稳定性,正成为药物发现中极具潜力的分子骨架。然而,将生成模型从线性肽扩展至环肽设计仍具挑战性,因为环化通过耦合的几何和生物物理约束大幅限制了可行的设计空间。此外,有限的训练数据使得现有方法大多依赖零样本生成或事后过滤,导致可行设计的产率较低,且对多目标权衡的控制有限。为解决这些局限,我们提出FAR-DPO(Feasibility-Aware and Robust Direct Preference Optimization,可行性感知与鲁棒直接偏好优化),这是一种与架构无关的框架,可引导生成模型生成在结构和生物物理上可行的环肽,尤其针对具有挑战性的靶点。FAR-DPO将可行性感知偏好构建与难度感知的分组鲁棒优化相结合,具体而言,它通过可行性门控多目标优势构建靶点内偏好对,并根据当前偏好损失自适应地重新加权预定义的难度组。在CPSea LNR基准测试中,在固定生成预算下,FAR-DPO将PepGLAD的总体成功率从46.89%提升至57.79%,将PepFlow的总体成功率从47.96%提升至49.57%。这些提升也延伸至最难的靶点四分位数,同时伴随更优的单靶点最佳结合分数。综上,这些结果证明了FAR-DPO在提升可行性和靶点鲁棒性方面的有效性。
英文摘要:
Cyclic peptides are emerging as promising molecular scaffolds in drug discovery due to their high binding affinity and structural stability. However, extending generative models from linear to cyclic peptide design remains challenging, as cyclization sharply restricts the feasible design space through coupled geometric and biophysical constraints. Moreover, limited training data has led existing approaches to rely largely on zero-shot generation or post hoc filtering, resulting in low yields of feasible designs and limited control over multi-objective trade-offs. To address these limitations, we propose FAR-DPO (Feasibility-Aware and Robust Direct Preference Optimization), an architecture-agnostic framework that steers generative models toward structurally and biophysically feasible cyclic peptide designs, particularly for challenging targets. FAR-DPO integrates feasibility-aware preference construction with difficulty-aware group-robust optimization. Specifically, it constructs within-target preference pairs through feasibility-gated multi-objective dominance and adaptively reweights predefined difficulty groups according to their current preference losses. On the CPSea LNR benchmark, under a fixed generation budget, FAR-DPO increases overall success rate from 46.89% to 57.79% on PepGLAD and from 47.96% to 49.57% on PepFlow. These gains also extend to the hardest target quartile and are accompanied by more favorable best-per-target binding scores. Together, these results demonstrate FAR-DPO's effectiveness in improving feasibility and target-wise robustness.