arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于单纯形精确嵌入的学术主题概貌的分类感知距离

Taxonomy-aware distances between scholarly topic profiles via an exact simplex embedding

Dmitry Gubanov, Alexander Chkhartishvili

arXiv 2608.22546首次发表:更新:

AI 中文总结

该研究针对学术主题概貌的距离计算问题,提出一种基于带权分类法的单纯形嵌入算子,实现了考虑分类层次的非退化度量,在OpenAlex分类法上验证了其有效性。

AI 中文摘要

主题概貌将出版物、作者及其他学术实体表示为固定主题集合上的概率分布,但平面总变差将每对不同的纯主题概貌视为最大分离,因此忽略了分类学邻近性。从带权有根分类法中,我们推导了一个基数归一化线性算子,该算子将叶主题映射到原概率单纯形中的点,并在总变差下精确实现归一化最低共同祖先超度量。该算子是双随机且正定的;在非负边聚类Gram算子类中,其归一化在简化分支树上唯一确定。对任意主题混合应用同一可逆算子,可得到非退化的层次感知度量,该度量收缩了平面总变差,与混合上的树Wasserstein距离不同,且可在O(|V|+L)的时间和内存中评估,无需形成稠密矩阵。在包含4516个终端主题的冻结OpenAlex分类法中,主题文本间的原始不相似度与分类学邻近性表现出一致的序数对齐,而253个校准内部节点中仅需3个进行单调校正。不过编码器选择会影响单个高度估计。该框架精确实现了提供的带权层次结构;文本仅用于初始化其节点高度,随后在诱导几何中计算学术主题概貌间的距离。

英文摘要

Topic profiles represent publications, authors, and other scholarly entities as probability distributions over a fixed set of topics, but flat total variation treats every pair of distinct pure-topic profiles as maximally separated and therefore ignores taxonomic proximity. From a rooted weighted taxonomy, we derive a cardinality-normalized linear operator that maps the leaf topics to points in the original probability simplex and exactly realizes a normalized lowest-common-ancestor ultrametric under total variation. The operator is doubly stochastic and positive definite; within the class of nonnegative edge-cluster Gram operators, its normalization is uniquely determined on the reduced branching tree. Applying the same invertible operator to arbitrary topic mixtures yields a nondegenerate hierarchy-aware metric that contracts flat total variation, differs from the tree-Wasserstein distance on mixtures, and can be evaluated in O(|V|+L) time and memory without forming the dense matrix. In a frozen OpenAlex taxonomy with 4,516 terminal Topics, raw dissimilarities between Topic texts showed consistent ordinal alignment with taxonomic proximity, while only 3 of 253 calibrated internal nodes required monotonic correction. Encoder choice nevertheless affected individual height estimates. The framework exactly realizes a supplied weighted hierarchy; text is used only to initialize its node heights, and distances between scholarly topic profiles are then computed in the induced geometry.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑