arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SciTBERT:用于科技语言处理的时间一致性语言模型家族

SciTBERT: A family of chronologically consistent language models for scientific and technological language processing

Thomas Gebhart, Russell J. Funk

arXiv 2610.12207首次发表:更新:

发表机构

Carlson School of Management, University of Minnesota(明尼苏达大学卡尔森管理学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出SciTBERT及SciTBERT-CI时间一致性语言模型家族,通过时间约束训练优化科技领域表示,在PatRepEval基准等任务上表现优于前代领域特定模型。

AI 中文摘要

预训练Transformer模型正越来越多地被用于研究科技进步。针对论文或专利文本调优的编码器,在科技领域内的下游分类、回归和相似度任务上,表现优于通用模型。然而,由于这些预训练模型固有的前瞻偏差和领域偏差,它们在研究科学、技术及其交叉领域的时间依赖特性或档案属性时,适用性受到限制。这些限制源于模型在训练时使用了时间顺序和文本源分布不受约束的语料库。我们提出SciTBERT:一系列时间一致性的BERT衍生语言模型,其训练数据来自科学论文、专利和高质量教育网页文本,训练数据的截止日期覆盖2013年至2025年的每一年。我们还利用论文和专利引用,以时间一致性的方式对这些模型进行后训练,创建了SciTBERT-CI模型家族。我们发现,即使在语料库受限于早期年份的训练数据时,这些模型通常也优于前代领域特定编码器模型。为进一步探究这类模型学习能弥合科学与技术鸿沟的表示的程度,我们推出了PatRepEval基准,这是一套针对科学-技术交叉领域的专利相关文本嵌入任务。在涵盖论文和专利的多种分类、回归和检索任务上的性能,凸显了使编码器模型表示与其下游任务的领域分布保持一致的重要性,且时间一致性编码器的表现可与未受时间约束训练的模型相媲美或更优。

英文摘要

Pre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology. However, the applicability of these models for studying time-dependent or archival properties of science, technology, and their interface is limited due to lookahead and domain biases inherent to these pre-trained models. These limitations arise from training on corpora with unconstrained chronological and text source distributions. We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025. We also post-train these models in a chronologically-consistent manner using paper and patent citations, creating SciTBERT-CI model family. We find that these models generally outperform predecessor domain-specific encoder models even when training data is limited by early year restrictions in the corpus. To further investigate the extent to which this class of models can learn representations that bridge science and technology, we introduce the PatRepEval benchmark, a suite of patent-related text embedding tasks at the science-technology interface. Performance in a variety of classification, regression, and retrieval tasks spanning papers and patents highlights the importance of aligning encoder model representations with the domain distributions of their downstream tasks, and chronologically consistent encoders can match or exceed models trained without temporal constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑