分子表示塑造基于流的聚合物生成中目标保真度与探索之间的平衡
Molecular representation shapes the balance between target fidelity and exploration in flow based polymer generation
浏览论文内容
中文总结 AI 辅助
本文提出PolyLatentFlow和LlamaUni,通过潜空间流匹配实现聚合物生成,在保持高有效性和低重放的同时,实现目标性能控制,并揭示分子表示对探索与保真度平衡的关键作用。
中文摘要 AI 辅助
设计具有目标性能的聚合物需要在有限的标记数据中探索广阔的化学空间。本文介绍了PolyLatentFlow,一个基于潜空间中连续时间流匹配的框架,用于无条件和条件聚合物生成,以及LlamaUni,一种结合聚合物序列和三维结构信息的多模态表示。在无条件生成中,带有LlamaUni的PolyLatentFlow在评估的无条件生成器中产生了相对于PolyInfo最大的新颖有效候选物产率,同时保持了高多样性。对于$T_g$条件生成,生成的性质分布在200°C目标范围内系统性偏移。在多性质任务中,分子表示在替代目标保真度上表现相似,但在有效性、训练集重放以及与标记聚合物的结构接近度上存在显著差异。带有LlamaUni的PolyLatentFlow始终结合了高有效性和低重放,并在CO$_2$/N$_2$条件生成中实现了每次尝试非重放目标命中的最大产率。这些结果展示了用于聚合物逆向设计的潜空间流匹配,并确定分子表示为目标控制和超越标记化学探索的关键决定因素。
英文摘要
Designing polymers with targeted properties requires navigating vast chemical spaces from limited labeled data. Here we introduce PolyLatentFlow, a framework based on continuous-time flow matching in latent space for unconditional and conditional polymer generation, together with LlamaUni, a multimodal representation combining polymer sequence and 3D structural information. In unconditional generation, PolyLatentFlow with LlamaUni produced the largest yield of valid candidates novel relative to PolyInfo among the evaluated unconditional generators while maintaining high diversity. For $T_g$ conditioning, generated property distributions shifted systematically across a 200 °C target range. In multi-property tasks, molecular representations showed similar surrogate target fidelity but differed markedly in validity, training-set replay, and structural proximity to labeled polymers. PolyLatentFlow with LlamaUni consistently combined high validity with low replay and achieved the largest per-attempt yield of nonreplayed target hits for CO$_2$/N$_2$ conditioning. These results demonstrate latent space flow matching for polymer inverse design and identify molecular representation as a key determinant of target control and exploration beyond labeled chemistry.
发表机构
- University of Delaware(特拉华大学)
机构由 AI 辅助整理,请以论文原文为准。