发表机构
IPRally Technologies Oy(IPRally科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出神经解析器,用双仿射注意力将专利文本直接转为发明图,复杂度降为线性,蒸馏自百万规则文档,以更低成本提升下游检索的引用召回率。
AI 中文摘要
专利检索需要处理通常超过数万词元的文档。大多数神经检索方法在截断的输入上运行,限制了其有效性。基于图的检索通过将每项专利表示为结构化的发明图来解决这一问题,但构建这些图依赖于脆弱的基于规则的解析器。我们提出了神经解析器,它改编了依存解析中的双仿射注意力,直接从专利文本预测发明图。我们的局部双仿射注意力将成对评分限制在滑动窗口内,将复杂度从$O(n^2)$降低到$O(n \cdot w)$。由于局部和全局评分共享相同的权重,该模型在短序列上训练,并在超过40,000词元的文档上部署,无需重新训练。从100万篇规则解析文档中蒸馏而来,它以3倍更低的推理成本超越了其教师模型:在下游图变换器检索系统中,神经图在短查询上将引用召回率提高了0.5%,在完整文档上提高了1.1%。
英文摘要
Patent search requires processing documents routinely exceeding tens of thousands of tokens. Most neural retrieval approaches operate on truncated inputs, limiting their effectiveness. Graph-based retrieval addresses this by representing each patent as a structured invention graph, but constructing these graphs relies on brittle rule-based parsers. We present the neural parser, which adapts biaffine attention from dependency parsing to predict invention graphs directly from patent text. Our local biaffine attention restricts pairwise scoring to a sliding window, reducing complexity from $O(n^2)$ to $O(n \cdot w)$. Since local and global scoring share the same weights, the model trains on short sequences and deploys on documents exceeding 40,000 tokens without retraining. Distilled from 1 million rule-parsed documents, it surpasses its teacher at 3$\times$ lower inference cost: neural graphs improve citation recall by 0.5% on short queries and 1.1% on full documents in a downstream Graph Transformer retrieval system.
CommentsAccepted for publication at the ECML PKDD 2026 conference (Applied Data Science track)