arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从涌现期信号学习与预测专利技术重用轨迹

Learning and Predicting Patent Technology Reuse Trajectories from Emergence-Time Signals

Ayham Yousef, Qiang Ye, Qiang Cheng

arXiv 2610.00806首次发表:更新:

AI 中文总结

本研究利用GRU自编码器与k-means从专利重用轨迹构建标签,基于七个涌现期特征以0.914的ROC-AUC预测重用模式,并揭示涌现期特征分群不等同于重用分类法。

AI 中文摘要

预测一项新出现的专利技术将如何被重用是技术情报的核心问题,但重用模式标签并不预先存在:它们必须从轨迹本身构建,而构建方式决定了预测的含义。我们研究了201,710项新型专利技术(首次出现的IPC代码配对,USPTO 2002-2022年)。我们的主要标注方法是在一个基于20年重用轨迹训练的GRU自编码器的潜在空间中应用k-means聚类,不使用任何手工设计的特征;据我们所知,这是首次将学习到的序列表示用于此任务。七个在技术出现第一年即可观测的涌现期特征,以一对多宏平均ROC-AUC为0.914恢复了这些标签,但仅使用涌现的日历年即可达到0.874。对观测满十年且无右删失的技术进行的重复实验表明,这一日历年效应主要反映了随时间变化的专利内容。我们另外在分形自编码器特征选择后,对涌现期特征本身进行聚类。该涌现期特征分群与基于GRU的标签仅略高于随机水平的一致性(调整兰德指数约0.04),因此涌现期特征的分群并非重用模式分类法,不应被解读为分类法。在遵循已发表构建的轨迹形状任务上,GBDT使用全部七个特征达到0.831,使用选定的特征子集达到0.740;已发表的0.728来自不同语料库和标注,是参考点而非基准。两个特征对于最早队列还存在左截断,我们对此进行了量化。

英文摘要

Forecasting how a newly emerged patent technology will be reused is central to technology intelligence, but reuse-pattern labels do not exist in advance: they must be constructed from the trajectories themselves, and how they are constructed determines what a forecast means. We study $201{,}710$ novel patent technologies (first-time IPC code pairings, USPTO 2002--2022). Our primary labeling applies $k$-means in the latent space of a GRU autoencoder trained on the $20$-year reuse trajectories, using no hand-crafted features; to our knowledge this is the first use of a learned sequence representation for this task. Seven emergence-time features, observable in a technology's first year, recover these labels at a one-vs-rest macro ROC-AUC of $0.914$, but the calendar year of emergence alone reaches $0.874$. A replication on technologies observed for ten full years, none of them right-censored, indicates that this calendar-year effect mainly reflects change over time in what was patented. Separately we cluster the emergence-time features themselves, after Fractal Autoencoder feature selection. That emergence-profile partition agrees with the GRU-based labels only marginally above chance (Adjusted Rand Index $\approx 0.04$), so a partition of emergence-time features is not a reuse-pattern taxonomy and should not be read as one. On a trajectory-shape task following the published construction, GBDT reaches $0.831$ with all seven features and $0.740$ with the selected subset; the published $0.728$, from a different corpus and labeling, is a reference point rather than a benchmark. Two features are additionally left-truncated for the earliest cohorts, which we quantify.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑