arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语法与语义融合:理解科学公式

Syntax Meets Semantics: Understanding Scientific Formulae

Yuni Susanti, Moritz Schubotz

arXiv 2608.02457首次发表:更新:

发表机构

FIZ Karlsruhe(FIZ卡尔斯鲁厄)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对科学公式语法与语义表征不匹配问题,采用图编码器、文本编码器结合对比学习实现跨模态对齐,显著提升了学术信息检索中的跨模态检索性能。

AI 中文摘要

科学公式是学术交流的基本组成部分,但其兼具结构化语法与语义载体的双重属性在学术信息检索领域尚未得到充分探索。尽管先前研究表明联合建模语法与语义模态可提升检索性能,但二者底层表征间的关系尚未得到系统研究。本研究实证探究公式语法与语义间的跨模态对应关系,发现尽管二者存在强潜在相关性,其原生表征空间却表现出极弱的可观测对应,表明两种模态间存在显著的表征不匹配。进一步评估是否可通过标准表征学习与对齐技术减少该不匹配:采用基于图的编码器表征语法结构,基于文本的编码器表征语义信息,再应用对比学习诱导共享表征空间。结果显示,学习到的对齐方式显著提升了跨模态检索性能,表明显式表征学习可恢复原始表征空间中缺失的对应关系。

英文摘要

Scientific formulae are a fundamental component of scholarly communication, yet their dual nature -- as structured syntax and carriers of semantics -- remains underexplored in scholarly information retrieval. Although prior studies show that jointly modeling syntactic and semantic modalities improves retrieval performance, the relationship between their underlying representations has not been systematically investigated. In this work, we empirically study cross-modal correspondence between formula syntax and semantics. We find that their native representation spaces exhibit extremely weak observable correspondence despite strong latent correlation, indicating a substantial representation mismatch between the two modalities. We further evaluate whether this mismatch can be reduced using standard representation learning and alignment techniques. We represent syntactic structure using graph-based encoders and semantic information using text-based encoders, then apply contrastive learning to induce a shared representation space. Results show that the learned alignment substantially improves cross-modal retrieval, suggesting that explicit representation learning can recover correspondence absent from the original representation spaces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑