发表机构
Asia School of Business(亚洲商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出超球面语义轨迹分析(HSTA),利用无监督方法从学术预印本和专利文本中追踪技术扩散,通过语义漂移和商业化偏移指标揭示大语言模型与AI系统领域的快速演化,并强调算力约束对文本信号的重要性。
AI 中文摘要
宏观经济生产率指标(如全要素生产率)由于行政调查间隔和国民核算惯例,在记录技术突破时存在多年的报告滞后。本文提出超球面语义轨迹分析(HSTA),一种无监督的定量方法,直接从非结构化的科学和商业文本流中追踪技术扩散。我们分析了来自arXiv学术预印本和美国专利商标局(USPTO)专利申请记录的30,000条过滤文档记录。通过使用球形K均值聚类在八个主要子主题上进行投影,将高维Transformer句子嵌入映射到单位超球面,并结合UMAP流形降维,HSTA形式化了两个定量指标:(1)语义质心向量漂移,用于追踪时间子语料库之间的词汇变化以识别结构范式转变;(2)商业化偏移,用于评估科学发现与知识产权申请之间跨语料库峰值密度对齐。将季度主题卷速度与来自Epoch AI数据库的物理硬件指标相关联,向量自回归F检验表明,仅季度论文卷速度在常规显著性水平下并不能格兰杰导致前沿算力分配激增,这凸显了将文本信号以物理资本约束为条件的必要性。实证结果显示,涵盖大语言模型(漂移指标为0.332)和人工智能系统(漂移指标为0.234)的子主题经历最高的语义演化速率,为补充传统经济统计提供了一种客观、实时的机制。
英文摘要
Macroeconomic productivity metrics, such as Total Factor Productivity, register technological breakthroughs with multi-year reporting lags due to administrative survey intervals and national accounting conventions. This paper introduces Hyperspherical Semantic Trajectory Analysis (HSTA), an unsupervised quantitative methodology that tracks technology diffusion directly from unstructured scientific and commercial text streams. We analyze 30,000 filtered document records spanning academic preprints from arXiv and patent application records from the USPTO. By projecting high-dimensional Transformer sentence embeddings onto unit hyperspheres using Spherical K-Means clustering across eight primary sub-topics and UMAP manifold reductions, HSTA formalizes two quantitative metrics: (1) Semantic Centroid Vector Drift, which tracks vocabulary shifts between temporal sub-corpora to identify structural paradigm transformations; and (2) Commercialization Offset, which evaluates cross-corpus peak density alignments between scientific discovery and intellectual property filings. Linking quarterly topic volume velocity with physical hardware metrics from the Epoch AI database, Vector Autoregressive F-tests demonstrate that quarterly paper volume velocity alone does not Granger-cause frontier compute allocation surges at conventional statistical significance levels, highlighting the necessity of conditioning textual signals on physical capital constraints. Empirical results reveal that sub-topics covering Large Language Models (with a drift metric of 0.332) and Artificial Intelligence Systems (with a drift metric of 0.234) undergo the highest rate of semantic evolution, offering an objective, real-time mechanism to complement traditional economic statistics.