arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26171cs.IRmath.PR

当隐藏链接无法被恢复:离岸泄露网络中的结构可辨识性界限与评估陷阱

When Concealed Links Cannot Be Recovered: A Structural Identifiability Bound and Evaluation Pitfalls in Offshore Leak Networks

Joseph Bingham

首次发表
浏览论文内容

中文总结 AI 辅助

本研究证明离岸泄露网络中隐藏链接大多无法通过链接预测恢复,提出无分布可辨识性界限,并揭示五个评估陷阱,建议将研究重点转向隐藏机制。

中文摘要 AI 辅助

巴拿马文件和天堂文件等泄露事件暴露了庞大的离岸实体网络,并引出了一个网络学习中的明显问题:这些结构旨在隐藏的关系——尤其是谁受益拥有什么——能否通过链接预测从泄露的公开部分中恢复?我们认为答案大多是否定的,而那些暗示相反的评估是在衡量错误的对象。我们的主要结果是一个无分布的可辨识性界限。对于任何尊重图同构且涵盖所有拓扑链接预测评分以及所有消息传递图神经网络的恢复规则,在观测图中被孤立隐藏的端点与其结构孪生体不可区分,因此其隐藏边无法以高于随机水平的概率被恢复。孤立只是最尖锐的情况。一般而言,恢复的上限由节点的结构不可区分性类别的大小决定,其中Weisfeiler-Leman颜色类是可计算的替代,而节点度数最多只是一个粗略的代理。在完整的ICIJ离岸泄露图(81.4万个实体和8.4万条标记的受益所有者边)上,一个无分类器、受度数控制的探针精确复现了孤立所有者(占所有所有者的20.6%)的0.5下限,而训练过的图神经网络也落在同一下限上。恢复率仅随着结构独特性的增长而上升,且该下限在五个泄露中的每一个中重新出现。在此过程中,我们记录了五个评估陷阱。每一个都使一个无法被击败的界限看起来被击败了,我们给出了一个避免它们的简短规则。实际意义是将努力从几乎不可恢复的隐藏委托人转向隐藏机制,并阐明为什么融合外部数据的效果远低于人们的期望。

英文摘要

Leaks such as the Panama and Paradise Papers expose large networks of offshore entities, and they invite an obvious question for network learning. Can the relations these structures are built to hide---above all, who beneficially owns what---be recovered from the public part of the leak by link prediction? We argue that the answer is mostly no, and that the analyses which suggest otherwise are measuring the wrong thing. Our main result is a distribution-free identifiability bound. For any recovery rule that respects graph isomorphism, and that covers every topological link-prediction score together with every message-passing graph neural network, a concealed endpoint left isolated in the observed graph is interchangeable with its structural twins, so its hidden edge cannot be recovered above chance. Isolation is only the sharpest case. In general the ceiling on recovery is set by the size of a node's structural-indistinguishability class, for which the Weisfeiler--Leman colour class is a computable stand-in, and node degree is at best a loose proxy. On the full ICIJ Offshore Leaks graph (814K entities and 84K labelled beneficial-owner edges) a classifier-free, degree-controlled probe reproduces an exact $0.5$ floor for isolated owners, who make up $20.6\%$ of all owners, and a trained graph neural network lands on the same floor. Recovery climbs only as structural distinctiveness grows, and the floor reappears in every one of the five leaks. Along the way we document five evaluation traps. Each one makes a bound that cannot be beaten look beaten, and we give a short rule that avoids them. The practical upshot is to redirect effort from the hidden principal, which is close to unrecoverable, toward the machinery of concealment, and to spell out why fusing external data helps far less than one would hope.

发表机构

  • Technion – Israel Institute of Technology(以色列理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑