ETH-TraceBench:面向时间、协议和合约转移的以太坊DeFi大规模事件流基准
ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift
浏览论文内容
中文总结 AI 辅助
ETH-TraceBench提出以太坊DeFi事件流基准,评估时间、协议和合约转移下的表示学习,发现简单模型在聚合测试中表现优异但协议新颖性下性能骤降,强调困难迁移与受控输入作为评估核心。
中文摘要 AI 辅助
以太坊去中心化金融(DeFi)提供了交易级事件流的公开、带时间戳的记录,但相同的公开符号可能造成强大的机器学习捷径。我们推出ETH-TraceBench,一个在时间、协议、池/基础设施和符号转移下评估以太坊DeFi表示学习的基准。原始事件宇宙覆盖2021年1月至2025年12月,包含13.5亿笔带日志的交易和50.1亿条原始日志行。模型评估使用固定的911,267实例监督样本,在2021-2024年上训练,在2025年上半年上选择模型,并在2025年下半年上测试。简单模型在聚合时间测试上表现强劲:TraceStats-GB在规范DEX测试集上达到0.953宏F1,TopicEmitterHashMLP达到0.959。在协议新颖性下性能急剧下降,TraceStats-GB、TopicEmitterTrace-SGD和TopicEmitterHashMLP的宏F1分别为0.794、0.743和0.766,而严格未见池得分仍为0.927、0.897和0.935。Uniswap v4和Ekubo v1(两者均未出现在监督训练中)比完整测试集难度大得多。联合掩蔽发射器和主题标识将TopicEmitterTrace-SGD的DEX宏F1降至0.916,清算宏F1降至0.774。在日志索引排序事件上的标准Transformer相对于相同事件的确定性洗牌没有提供一致优势,表明高聚合得分可以在没有复杂时间建模的情况下出现。一项自然流行度审计估计2025年下半年DEX在以太坊日志交易中的流行度约为22.5%,一项确定性400笔交易审计发现与任务标签来源和独立重新查询的原始日志计数完全一致。因此,ETH-TraceBench将困难迁移和受控输入条件(而非单一聚合得分)作为主要评估目标。
英文摘要
Ethereum decentralized finance (DeFi) provides a public, time-stamped record of transaction-level event streams, but the same public symbols can create strong machine-learning shortcuts. We introduce ETH-TraceBench, a benchmark for evaluating Ethereum DeFi representations under temporal, protocol, pool/infrastructure, and symbolic shift. The raw event universe covers January 2021-December 2025 and contains 1.35 billion transactions with logs and 5.01 billion raw log rows. Model evaluation uses a fixed 911,267-instance supervised sample, training on 2021-2024, selecting models on 2025H1, and testing on 2025H2. Simple models perform strongly on the aggregate temporal test: TraceStats-GB reaches 0.953 macro-F1 and TopicEmitterHashMLP 0.959 on the canonical DEX test set. Performance drops sharply under protocol novelty, with macro-F1 of 0.794, 0.743, and 0.766 for TraceStats-GB, TopicEmitterTrace-SGD, and TopicEmitterHashMLP, while strict unseen-pool scores remain 0.927, 0.897, and 0.935. Uniswap v4 and Ekubo v1, both absent from supervised training, are materially harder than the full test. Jointly masking emitter and topic identity reduces DEX macro-F1 to 0.916 and liquidation macro-F1 to 0.774 for TopicEmitterTrace-SGD. A standard Transformer over log-index-ordered events provides no consistent advantage over a deterministic shuffle of the same events, indicating that high aggregate scores can arise without sophisticated chronological modeling. A natural-prevalence audit estimates 2025H2 DEX prevalence among logged Ethereum transactions at about 22.5%, and a deterministic 400-transaction audit finds complete agreement with task label sources and independently re-queried raw-log counts. ETH-TraceBench therefore treats difficult transfer and controlled-input conditions, rather than a single aggregate score, as the main evaluation target.
发表机构
- University College London(伦敦大学学院)
- University of Warwick(华威大学)
机构由 AI 辅助整理,请以论文原文为准。