发表机构
School of Computer Science and Technology, Zhejiang University; ZJU-UIUC Institute, Zhejiang University; Shaoxing K3i Technology Co. Ltd; State Key Laboratory of CAD&CG, Zhejiang University(浙江大学计算机科学与技术学院; 浙江大学ZJU-UIUC学院; 绍兴K3i科技有限公司; 浙江大学CAD&CG国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对专利匹配的LLM方法存在成本高、易遗忘等问题,本文提出自知识RAG框架,结合FAISS检索与生成式匹配机制,在真实专利数据集上提升了检索匹配准确率。
AI 中文摘要
基于大语言模型(LLM)的专利检索与匹配在知识产权保护中发挥着至关重要的作用。然而,由于专利文档结构复杂、技术术语密集且包含多模态信息,传统方法难以准确识别专利间的细微差异。现有的基于LLM的专利匹配方法通常依赖于领域特定的预训练或指令微调,这往往需要高昂的人工标注成本且易出现灾难性遗忘。尽管检索增强生成(RAG)方法引入了外部知识,但它们未能充分利用LLM自动解析专利和挖掘深层语义关系的能力。为解决这些局限,本文提出一种自知识RAG框架,该框架引导LLM从专利匹配查询中自主提取关键技术实体并构建分层本体结构,从而实现查询扩展与精准检索。该方法将FAISS检索与生成式匹配机制相结合,利用自知识增强模型对专利创新的理解,显著提升检索与匹配准确率。实验结果表明,所提方法在真实世界专利数据集上展现出优异性能,验证了其有效性与应用潜力。
英文摘要
Patent retrieval and matching based on large language models (LLMs) play a vital role in intellectual property protection. However, due to the complex structure of patent documents, dense technical terminology, and multi-modal information, traditional methods struggle to accurately identify subtle differences between patents. Existing LLM-based patent matching approaches typically rely on domain-specific pretrained or instruction tuning, which often entail high manual labeling costs and catastrophic forgetting. While retrieval-augmented generation (RAG) methods introduce external knowledge they fail to fully leverage LLM's capability to automatically parse patents and mine deep semantic relationships. To address these limitations, this paper proposes a self-knowledge RAG framework that guides LLMs to autonomously extract key technical entities and construct hierarchical ontological structures from patent matching queries, thereby enabling query expansion and precise retrieval. The method integrates the FAISS retrieval with a generative matching mechanism, leveraging self-knowledge to enhance the model's understanding of patent innovations and significantly improve retrieval and matching accuracy. Experimental results demonstrate the outstanding performance of the proposed method on real-world patent datasets, validating its effectiveness and application potential.
CommentsAccepted by IEEE CSCWD 2026
DOI:10.1109/CSCWD68734.2026.11582454