arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Omega-N:可解释的结构节点描述符及其适用域

Omega-N: Interpretable Structural Node Descriptors and Their Applicability Domain

Alberto Acedo

arXiv 2609.01633首次发表:更新:

发表机构

Biome Makers Inc(Biome Makers公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建了仅需图信息的10个可解释结构节点特征Omega-N,在节点分类等任务中表现优于或持平于递归特征引擎,在药物-靶点优先排序任务中提升了AUPRC,且结果可复现

AI 中文摘要

复合结构索引用单个数字概括网络;对于基于三角形的索引,其具有谱冗余性:Tr(A³)是邻接谱的三阶矩。非冗余内容位于更低层级,即diag(A³),其依赖于特征向量,无法由谱确定。针对该索引族的理论论文中的一个推论指出了这一点,并预测:全局标量应与锐化的谱基线关联而非超越它们,而节点级归因在结构震中数量未知时表现更好。本文对此进行了验证。我们通过定位四个因子各自的方式构建Omega-N。直接定位的条件很差;来自已发表实践的两种修正方案解决了该问题,即每个局部因子的配置-零过剩以及多个尺度下的个性化PageRank邻域,仅从图中即可为每个节点提供10个可解释特征,无需属性、训练或嵌入。与5级递归的递归特征引擎相比,在6个域内节点分类评估中,Omega-N在1个评估中胜出,在4个评估中持平,其特征数为10,而对方最多达252个。从图和标签(而非性能)计算出的两个统计量可无误差地划分8个基准,它们排除的两个基准正是Omega-N表现不佳的两个。最强应用是蛋白质相互作用网络上的药物-靶点优先排序:在4种构建方式下,相较于中心性组,AUPRC提升了0.073至0.144,在独立的AP-MS网络和标签源上得到重复,且通过了3种偏差控制(度匹配,10次重复:提升0.1047和0.1030,均为10/10,p=0.00195)。最明确的负面结果来自同一应用:将Omega-N添加到中心性特征加Node2Vec特征中无变化(提升0.0014,p=0.31)。该主张范围较窄:10个可解释特征

英文摘要

A composite structural index summarises a network in one number, and for a triangle-based index it is spectrally redundant: Tr(A^3) is the third moment of the adjacency spectrum. The non-redundant content sits one level down, in diag(A^3), which depends on eigenvectors and is not spectrally determined. A corollary in the theory paper predicted that the global scalar should tie sharpened spectral baselines rather than beat them, while the node-wise attribution should do better where the number of structural epicentres is unknown. We construct Omega-N by localizing each of the four factors. The direct localization is badly conditioned; two corrections from published practice fix it, a configuration-null excess per factor and a personalized-PageRank neighbourhood at several scales, giving ten interpretable features per node, with no attributes, training or embeddings. Against a recursive feature engine at five levels of recursion, Omega-N wins on three and ties on two of the six in-domain evaluations, the sixth a declared null where every arm returns chance, with ten features against its 28 to 252 before pruning. Two statistics from the graph and labels, not from performance, partition the eight benchmarks without error, and the two they exclude are the two on which it loses. The strongest application is drug-target prioritisation on protein interaction networks: +0.032 to +0.103 AUPRC over a six-feature centrality battery and +0.084 to +0.208 over the four-feature one, across three constructions, replicated on an independent AP-MS network and label source (degree-matched: +0.0723 on STRING, +0.0560 on BioPlex, p=0.00195). Adding Omega-N to centralities plus Node2Vec changes nothing. The claim is narrow and it is the point: ten named features, computed without training, match or beat hand-crafted centralities and a recursive engine, and do not touch learned representations.

Comments18 pages, 3 figures. Reference implementation, notebooks and data-preparation scripts at https://github.com/BiomeMakers/OmegaN

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑