一个开源的、事件驱动的加密货币市场数据处理流水线:数据采集、预测及链上欺诈检测
An Open-Source, Event-Driven Pipeline for Cryptocurrency Market Data: Ingestion, Forecasting, and On-Chain Fraud Detection
浏览论文内容
中文总结 AI 辅助
本文提出一款基于Apache Kafka等开源技术的事件驱动加密货币数据流水线,实现数据采集、存储,并用其数据对比ARIMA与LSTM的价格预测效果,结合分类器完成以太坊欺诈检测,同时明确相关评估局限性。
中文摘要 AI 辅助
加密货币市场会产生高频、多源的数据,除非团队已具备商用级的流处理和数据仓库基础设施,否则处理这些数据的成本很高。本文描述了一个完全开源的流水线,它在普通硬件上重现了云原生、事件驱动系统的行为——文件到达触发消息,消息触发计算,该流水线使用Apache Kafka和文件系统轮询器替代托管式云触发器。该流水线将Gemini交易所的历史数据划分为小时级和分钟级文件,通过两组独立分组的Kafka消费者异步采集这些数据(一组用于审计日志记录,另一组用于Spark触发的ETL),并将清洗后的输出存储到PostgreSQL数据仓库中,该仓库包含历史和聚合模式以及特定资产的数据集市。我们利用生成的比特币数据集市,将季节性ARIMA模型与单层LSTM网络进行价格预测对比,并结合额外的工程特征,将随机森林(Random Forest)和梯度提升(Gradient Boosting)分类器分别应用于Farrugia等人提出的公开以太坊欺诈检测基准。我们报告了该流水线的架构、建模方法及所得指标,同时明确指出了在不同时间范围发布的预测进行对比的局限性,以及在静态、已标注的数据集上评估欺诈检测的局限性。
英文摘要
Cryptocurrency markets generate high-frequency, multi-source data that is expensive to work with unless a team already has commercial-grade streaming and warehousing infrastructure in place. This paper describes a fully open-source pipeline that reproduces the behavior of a cloud-native, event-driven system -- file arrival triggering a message, a message triggering compute -- entirely on commodity hardware, using Apache Kafka and a filesystem-watching poller in place of managed cloud triggers. The pipeline partitions historical Gemini exchange data into hourly and minutely files, ingests them asynchronously through two independently grouped Kafka consumers (one for audit logging, one for Spark-triggered ETL), and lands cleaned output in a PostgreSQL warehouse with historical and aggregated schemas plus asset-specific data marts. We use the resulting Bitcoin data mart to compare a seasonal ARIMA model against a single-layer LSTM network for price forecasting, and separately apply Random Forest and Gradient Boosting classifiers, with additional engineered features, to the public Ethereum fraud detection benchmark introduced by Farrugia et al. We report the architecture, the modeling methodology, and the resulting metrics, and we are explicit about the limitations of comparing forecasts issued at different horizons and of evaluating fraud detection on a static, already-labeled dataset.