arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于反应产率预测的通用视觉与交叉注意力机制

Generic Vision and Cross-Attention for Reaction Yield Prediction

Qiwei Han, Chi Zhou

arXiv 2608.00776首次发表:更新:

发表机构

Duke University; Georgia Institute of Technology(杜克大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对传统反应产率预测缺乏空间信息的问题,提出双模态视觉交叉注意力架构,结合通用视觉骨干网络与残差连接,实现了优于传统方法的预测精度,为物理化学研究提供了可扩展的深度视觉学习方案。

AI 中文摘要

传统反应产率预测受限于缺乏显式空间信息的一维量子描述符。为解决该问题,本文提出双模态视觉交叉注意力架构,将表格形式的物理有机数据与二维分子拓扑结构相融合。值得注意的是,处理简单二维骨架结构的通用计算机视觉骨干网络,其性能优于纯量子基线方法。通过协同两种模态,最优交叉注意力框架实现了优于传统方法的预测精度(测试均方根误差=5.27%)。经机制探究发现,存在活跃的、由描述符引导的空间查询,可将宏观空间位阻识别有效转移至视觉通路;此外,网络会学习动态化学层级,重点关注芳基卤化物等关键空间位阻瓶颈。同时,采用残差跳跃连接保护非空间电子参数,避免其在融合过程中被破坏性衰减。总体而言,本文提供了一个可扩展且高可解释性的蓝图,用于通过深度视觉学习增强物理化学研究。

英文摘要

Traditional reaction yield prediction is constrained by 1D quantum descriptors that lack explicit spatial information. To address this gap, a dual-modal Vision Cross-Attention architecture is proposed, fusing tabular physical-organic data with 2D molecular topologies. Notably, it is demonstrated that a generic computer vision backbone processing simple 2D skeletal structures independently outperforms purely quantum-based baselines. By synergizing both modalities, superior predictive accuracy compared to traditional methodologies is achieved by the optimal cross-attention framework (Test RMSE = 5.27%). Through mechanistic probing, active, descriptor-guided spatial querying is observed, effectively offloading macroscopic steric identification to the visual pathway. Furthermore, a dynamic chemical hierarchy is learned by the network to heavily prioritize critical steric bottlenecks, such as the aryl halide. Concurrently, residual skip connections are utilized to protect non-spatial electronic parameters from destructive attenuation during fusion. Collectively, a scalable and highly interpretable blueprint is provided for augmenting physical chemistry with deep visual learning.

Comments12 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑