arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2402.19411cs.IRcs.CLcs.LG

PaECTER:使用引文感知Transformer的专利级表示学习

PaECTER: Patent-level Representation Learning using Citation-informed Transformers

  • Max Planck Institute for Innovation and Competition(马克斯·普朗克创新与竞争研究所)

机构由 AI 辅助整理,请以论文原文为准。

Mainak Ghosh, Michael E. Rose, Sebastian Erhardt, Erik Buunk, Dietmar Harhoff

更新

AI总结:

本研究提出专为专利设计的开源文档级编码器PaECTER,通过审查员引文信息微调BERT for Patents生成专利表示,在专利相似性任务上性能优于现有主流模型,可支撑现有技术检索等下游应用。

AI中文摘要:

PaECTER是一款专为专利设计的开源文档级编码器。我们利用审查员添加的引文信息对BERT for Patents进行微调,以生成专利文献的数值表示。PaECTER在相似性任务上的表现优于当前专利领域使用的最优模型。具体而言,在我们的专利引文预测测试数据集上,该模型在不同排名评估指标上的表现均优于专利专用预训练语言模型(BERT for Patents)以及通用文本嵌入模型(如E5、GTE和BGE)。当与25篇不相关专利进行对比时,PaECTER预测出至少一篇最相似专利的平均排名为1.32。PaECTER从专利文本生成的数值表示可用于分类、知识流追踪或语义相似性搜索等下游任务。语义相似性搜索在发明人和专利审查员的现有技术检索场景中尤为重要。

英文摘要:

PaECTER is an open-source document-level encoder specific for patents. We fine-tune BERT for Patents with examiner-added citation information to generate numerical representations for patent documents. PaECTER performs better in similarity tasks than current state-of-the-art models used in the patent domain. More specifically, our model outperforms the patent specific pre-trained language model (BERT for Patents) and general-purpose text embedding models (e.g., E5, GTE, and BGE) on our patent citation prediction test dataset on different rank evaluation metrics. PaECTER predicts at least one most similar patent at a rank of 1.32 on average when compared against 25 irrelevant patents. Numerical representations generated by PaECTER from patent text can be used for downstream tasks such as classification, tracing knowledge flows, or semantic similarity search. Semantic similarity search is especially relevant in the context of prior art search for both inventors and patent examiners.

补充信息

↑