arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向科学文献表示的文档内非对称预测学习

Asymmetric Within-Document Predictive Learning for Scientific Document Representation

You Zuo, Éric de la Clergerie, Benoît Sagot

arXiv 2608.28625首次发表:更新:

发表机构

Questel; Inria(奎斯特尔; 法国国家信息与自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出无引文框架SciJEPA,采用文档内非对称预测学习科学文献表示,结合SIGReg可提升性能,为科学文献表示提供了新的无引文方案。

AI 中文摘要

我们利用论文的语篇结构研究科学文献表示的预测式预训练,提出SciJEPA这一无引文框架,通过文档内非对称预测进行学习:使用标题和摘要的表示来预测方法部分的表示,再用方法部分的表示预测结论部分的表示。在RELISH、高影响力引文、SciDocs及引文预测任务上的实验表明,纯预测式训练可行,但弱于使用相同章节对的受控对比基线;加入切片各向同性高斯正则化(SIGReg)可大幅提升性能并缩小差距,该正则化效果具有任务依赖性:适度SIGReg有助于细粒度排序,较强正则化则会削弱局部对齐。我们进一步证实,不同编码分支支持不同检索机制。这些结果表明,只要对嵌入几何进行仔细控制,文档内预测学习可作为无引文的有效补充,用于科学文献表示。

英文摘要

We study predictive pretraining for scientific document representation using the discourse structure of papers. We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations. Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrastive baseline using the same section pairs. Adding Sliced Isotropic Gaussian Regularization (SIGReg) substantially improves performance and narrows this gap. The effect of regularization is task-dependent: moderate SIGReg helps fine-grained ranking, while stronger regularization can weaken local alignment. We further show that different encoding branches support different retrieval regimes. These results position within-document predictive learning as a promising citation-free complement for scientific document representation, provided that embedding geometry is carefully controlled.

Journal ref(ARTS)@TALN 2026 - Atelier ''Analyse et Recherche de Textes Scientifiques'', Jun 2026, Nantes, France

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑