AI 中文总结
研究针对X射线吸收光谱建模中异构模态学习难题,提出Uni-XAS框架。通过XASLip解决元素内配位变化,用检索增强解码防尺度崩溃等,还引入排列校正流匹配解决逆问题,在多方面表现出色,为多模态学习和评估奠定基础。
AI 中文摘要
X射线吸收光谱(XAS)是探测局部原子环境的关键技术,但基于学习的建模必须连接一维连续光谱和三维原子结构这两种异构模态。现有方法通常将正向光谱预测和反向结构推断解耦为单独的回归任务,阻碍了共享表示学习。此外,相同原子之间严重的排列模糊性通常将反向建模限制为粗结构描述符,而不是显式的三维结构生成。在这项工作中,我们提出了Uni-XAS,一个统一的基准和学习框架,将双向XAS建模重新构建为跨模态对齐和条件生成问题。我们首先提出XASLip,一种将物理感知光谱编码器与吸收器感知流形优化策略相结合的对齐方法,以解决细粒度的元素内配位变化。在此共享潜在空间的基础上,我们将正向预测公式化为通过检索增强解码进行锚定绝对光谱生成,有效防止物理尺度崩溃和能量漂移。对于本质上不适定的逆问题,我们引入了排列校正流匹配,将类型明智的最优传输集成到连续生成流中,以在不依赖于重型高阶等变架构的情况下为配体排列模糊性提供原则性解决方案。在32,8839个结构-光谱对的大规模标准化基准上进行评估,Uni-XAS在跨模态检索、准确的绝对光谱预测和成分条件三维结构生成方面表现出强大性能,为科学光谱中的多模态学习和标准化评估建立了一个可扩展、可重复和协议一致的基础。
英文摘要
X-ray absorption spectroscopy (XAS) is a key technique for probing local atomic environments, yet learning based modeling must bridge two heterogeneous modalities: 1D continuous spectra and 3D atomic structures. Existing approaches typically decouple forward spectrum prediction and inverse structure inference into separate regression tasks, hindering shared representation learning. Moreover, severe permutation ambiguity among identical atoms often limits inverse modeling to coarse structure descriptors rather than explicit 3D structure generation. In this work, we present Uni-XAS, a unified benchmark and learning framework that reframes bidirectional XAS modeling as a cross-modal alignment and conditional generation problem. We first propose XASLip, an alignment recipe coupling a physics-aware spectral encoder with an absorberaware manifold optimization strategy to resolve fine-grained intra-element coordination variations. Building upon this shared latent space, we formulate forward prediction as anchored absolute-spectrum generation via retrieval-augmented decoding, effectively preventing physical scale collapse and energy drift. For the inherently ill-posed inverse problem, we introduce Permutation-Rectified Flow Matching, which integrates type-wise optimal transport into a continuous generative flow to provide a principled solution to ligand permutation ambiguity without relying on heavy high-order equivariant architectures. Evaluated on a largescale standardized benchmark of 328,839 structure-spectrum pairs, Uni-XAS demonstrates strong performance in cross-modal retrieval, accurate absolute-spectrum prediction, and composition-conditional 3D structure generation, establishing a scalable, reproducible, and protocol-consistent foundation for multimodal learning and standardized evaluation in scientific spectroscopy.
Comments48 pages, 7 figures, 11 tables. Accepted by ACMMM 2026