质量重于数量:通过特征工程半监督检测非法比特币流动
Quality over Quantity: Semi-Supervised Detection of Illicit Bitcoin Flows via Feature Engineering
浏览论文内容
中文总结 AI 辅助
本研究提出半监督框架,利用高保真特征工程(如KeyLinker和SSU)在1.63亿笔交易中检测非法比特币流动,证明数据质量优于数量,F1达0.84。
中文摘要 AI 辅助
检测非法加密货币交易受到极端类别不平衡、对抗性混淆和可靠标签稀缺的阻碍。虽然半监督学习(SSL)通过利用未标记数据提供了一种有前景的解决方案,但我们表明其成功并非仅由数据量保证,而是取决于数据质量。我们引入了一个用于检测共享发送混合器(SSM)交易中非法比特币流动的SSL框架,该框架基于包含1.63亿笔交易的全面历史数据集。我们的主要结论是,SSL的成功取决于数据质量而非数据量:高保真特征,如KeyLinker地址聚类和共享发送解缠(SSU)复杂度指标,在未标记数据上实现了0.84的F1分数。最后,我们实证表明,常见启发式方法如一次性变更(OTC),尽管丰富,但会引入噪声,而战略性依赖更高保真特征如KeyLinker则至关重要。我们的工作确立了在区块链取证中,实现更好性能的途径在于更智能的特征工程以提升数据质量,而非仅仅依赖更大的数据集。
英文摘要
Detecting illicit cryptocurrency transactions is hampered by extreme class imbalance, adversarial obfuscation, and a scarcity of reliable labels. While semi-supervised learning (SSL) offers a promising solution by leveraging unlabeled data, we show that its success is not guaranteed by data volume alone but is contingent on data quality. We introduce an SSL framework for detecting illicit Bitcoin flows in Shared Send Mixers (SSM) transactions, built on a comprehensive historical dataset comprising 163 million transactions. Our main conclusion is that the success of SSL depends on data quality rather than volume: high-fidelity features such as KeyLinker address clustering and Shared Send Untangling (SSU) complexity metrics achieve an F1 score of 0.84 on unlabeled data. Finally, we empirically show that common heuristics like One-Time Change (OTC), though abundant, introduce noise, while strategic reliance on higher-fidelity features like KeyLinker is essential. Our work establishes that in blockchain forensics, the path to better performance lies in smarter feature engineering for data quality, not just larger datasets.
发表机构
- Skolkovo Institute of Science and Technology(斯科尔科沃科学技术研究院)
- Moscow Institute of Physics and Technology(莫斯科物理技术学院)
机构由 AI 辅助整理,请以论文原文为准。