arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

原子设计变换器:通过xTB奖励强化学习实现支架条件下的3D分子生成

Atomic Design Transformer: Scaffold-Conditioned 3D Molecule Generation with xTB-Verified Reinforcement Learning

Takao Kotani

arXiv 2607.15918首次发表:更新:

发表机构

Institute for Advanced Study, Kyoto University; Osaka University(京都大学高等研究院; 大阪大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出原子设计变换器ADT用于3D分子生成,通过自回归和令牌化实现SE(3)不变性,以xTB拓扑保留率评估。先预训练模型,再用RLVR强化学习提升性能,还提出逆运动学变换器,实现了直接3D分子生成。

AI 中文摘要

我们提出了一种用于3D分子生成的SE(3)不变变换器——原子设计变换器(ADT)。ADT通过自回归方式逐个放置原子,通过令牌化实现SE(3)不变性。其主干是普通因果变换器,令牌流完全指定3D结构及其化学键图G。为评估生成分子,引入xTB拓扑保留率(XTP)。我们评估了两个ADT模型,第一个在GEOM-Drugs≤30重原子数据集上预训练,第二个通过对抗可验证的xTB奖励的强化学习(RLVR)继续训练。RLVR将XTP提高到约98%,有效分子产率N^gen/N提高到约95%。最后,我们提出了一种逆运动学变换器,用于恢复大分子的XTP。ADT实现了直接3D生成。

英文摘要

We present an autoregressive 3D-molecule generator with SE(3)-invariant tokenization, the Atomic Design Transformer (ADT). ADT places atoms one at a time, autoregressively. SE(3) invariance is achieved by tokenization: each new atom's position is encoded in the local coordinate frame of a previously placed atom. The backbone is a plain causal transformer. The token stream fully specifies a 3D structure together with its chemical-bond graph G, without any bond-order assignment. The model emits heavy-atom skeletons; hydrogens are added by separate learned models before the xTB relaxation. To score generated molecules we introduce the xTB topology-preservation rate (XTP): the fraction of molecules for which an xTB GFN2 relaxation preserves G specified by the token stream. For XTP-accepted molecules we also report the relaxation energy and the root mean square of the atomic displacement (RMSD). We evaluate two ADT models. The first is ADT pretrained on the GEOM-Drugs $\le\!30$-heavy-atom dataset; we benchmark scaffold-conditioned 3D generation across seven drug-like scaffolds from the model. It reaches an XTP of ${\sim}55\%$ and a valid-molecule yield $N^{\mathrm{gen}}/N$ of ${\sim}53\%$, where $N^{\mathrm{gen}}/N$ is the fraction of samples that are distinct, topology-preserving, and RDKit-readable. The second model continues from the first by reinforcement learning against the verifiable xTB reward (RLVR), using no external molecules. RLVR raises XTP to ${\sim}98\%$ and $N^{\mathrm{gen}}/N$ to ${\sim}95\%$, while approximately preserving the GEOM-Drugs size and composition distributions. Finally, we present an Inverse-Kinematics Transformer that recovers XTP for large molecules, where discretization error accumulates. ADT thus enables direct 3D generation.

Comments8 pages. Code and data: github.com/tkotani/ADT (v2.1), doi:10.5281/zenodo.20635985; one instruction to an AI coding agent regenerates the benzene row of Table 2 from the released checkpoint (~1 h on one GPU). v2: new title; RLVR model released as rlvr_E240direct.pt; scaffold frame caches corrected; all generation numbers re-tabulated with one public script; text revised

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑