从地下论坛讨论中提取和验证非法比特币地址
Extracting and Verifying Illicit Bitcoin Addresses from Underground Forum Discussions
浏览论文内容
中文总结 AI 辅助
本研究提出结合LLM辅助筛选、专家审核与链上验证的可复现流程,从地下论坛HackForums构建含2438个经人工验证非法比特币地址的数据集,为相关研究提供支撑。
中文摘要 AI 辅助
现有带标注的比特币数据集大多来自社区举报的滥用情况、区块链启发式方法、特定事件收集或专有标注流程,其构建方法很少能公开复现,且往往只能提供有限的证据证明地址直接参与非法活动。我们提出了一种可复现的流程,用于从拥有十五年存档活动的地下网络犯罪论坛HackForums中构建有证据支持的比特币标注。该流程结合了大语言模型(LLM)辅助筛选、专家审核和链上验证,以识别论坛讨论中明确与非法交易相关联的比特币地址。每个发布的标注都有来自地下讨论的上下文证据支持,并经过链上验证。生成的数据集包含2010年至2024年间经人工验证的2438个非法比特币地址,以及在LLM筛选期间分配的十二个网络犯罪类别。我们发布该数据集、时间元数据和完整的提取流程,以支持针对加密货币便利化网络犯罪的可复现研究。
英文摘要
Existing labeled Bitcoin datasets are largely derived from community-reported abuse, blockchain heuristics, incident-specific collections, or proprietary labeling processes. Their construction methods are rarely publicly reproducible and often provide limited evidence that an address was directly involved in illicit activity. We present a reproducible pipeline for constructing evidence-backed Bitcoin labels from HackForums, an underground cybercrime forum with fifteen years of archived activity. The pipeline combines LLM-assisted screening, expert review, and on-chain validation to identify Bitcoin addresses explicitly associated with illicit transactions discussed on the forum. Each released label is supported by contextual evidence from underground discussions and validated on-chain. The resulting dataset contains 2,438 manually verified illicit Bitcoin addresses spanning 2010-2024 and twelve cybercrime categories assigned during LLM screening. We release the dataset, temporal metadata, and the complete extraction pipeline to support reproducible research on cryptocurrency-facilitated cybercrime.