arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GEqTrain:用于跨3D科学任务重新定位等变图神经网络的配置驱动框架

GEqTrain: A Configuration-Driven Framework for Retargeting Equivariant Graph Neural Networks Across 3D Scientific Tasks

Daniele Angioletti, Marco Nobile, Vittorio Limongelli

arXiv 2607.19083首次发表:更新:

发表机构

Faculty of Biomedical Sciences, Euler Institute, Universitá della Svizzera italiana(瑞士意大利语区大学 生物医学科学学院欧拉研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对等变图神经网络重用受限问题,提出配置驱动框架GEqTrain,可通过配置重新定位到新任务,在多3D科学任务展示灵活性,还引入GEqDiff,验证其能力,目标是提升等变建模的可重复性、扩展性和可重用性。

AI 中文摘要

等变图神经网络为三维科学数据提供了强大的建模语言,但其重用常受限于与特定任务、输出和训练机制相关的实现。我们提出了GEqTrain,这是一个配置驱动框架,它分离了数据集语义、模型组成和训练目标。原始数据被映射到类型化的节点、边和图级字段,而模型堆栈、损失和训练工作流程通过Hydra配置声明性地组装。一个共享的等变主干和训练基础设施因此可以主要通过配置重新定位到新任务。我们在一个软件堆栈中处理的三个不同问题上展示了这种灵活性:生物分子系统的粗粒度到原子级反向映射、分子固体中NMR化学位移的预测以及等变生成建模。我们的目标不是超越单独优化的特定任务系统,而是表明共享表示和训练基础设施可以通过配置更改在定性不同的任务中实现有竞争力的准确性。我们还引入了GEqDiff,一种基于等变流匹配的生成扩展。GEqDiff将用户定义的等变字段视为一等生成目标,在单个等变流中联合传输笛卡尔位置和跨越高达l = 3表示的非标量节点字段。我们在受蛋白质二级结构基序启发的受控合成基准上验证了这种能力,表明具有异构变换属性的字段可以联合且高保真地重建。通过减少在预测和生成、标量和张量设置之间切换的软件开销,GEqTrain旨在使等变建模更具可重复性、可扩展性和可重用性。

英文摘要

Equivariant graph neural networks provide a powerful modeling language for three-dimensional scientific data, but their reuse is often limited by implementations tied to specific tasks, outputs, and training regimes. We present GEqTrain, a configuration-driven framework that separates dataset semantics, model composition, and training objectives. Raw data are mapped to typed node-, edge-, and graph-level fields, while model stacks, losses, and training workflows are assembled declaratively through Hydra configurations. A shared equivariant backbone and training infrastructure can therefore be retargeted to a new task primarily through configuration. We demonstrate this flexibility on three different problems handled within one software stack: coarse-grained-to-atomistic backmapping of biomolecular systems, prediction of NMR chemical shifts in molecular solids, and equivariant generative modeling. Our aim is not to surpass individually optimized task-specific systems, but to show that a shared representation and training infrastructure can achieve competitive accuracy across qualitatively different tasks at the cost of a configuration change. We further introduce GEqDiff, a generative extension based on equivariant flow matching. GEqDiff treats user-defined equivariant fields as first-class generation targets, jointly transporting Cartesian positions and non-scalar node fields spanning representations up to l=3 within a single equivariant flow. We validate this capability on a controlled synthetic benchmark inspired by protein secondary-structure motifs, showing that fields with heterogeneous transformation properties can be reconstructed jointly and with high fidelity. By reducing the software overhead of moving between predictive and generative, scalar and tensorial settings, GEqTrain aims to make equivariant modeling more reproducible, extensible, and reusable.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑