发表机构
Zeru AI; Ohio State University(泽鲁人工智能公司; 俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对链上钱包声誉评分中手写公式不可微且无法随数据改进的问题,提出神经网络 zScore-N,以公式为教师生成训练数据,在保持高精度复现的同时显著降低数据缺失误差,并已投入生产。
AI 中文摘要
钱包声誉评分决定了谁获得空投、谁可以借贷、谁进入去中心化金融中的白名单。这些评分几乎总是以手写公式起步:对行为特征进行截断对数、线性和平方根变换的组合,每个阈值和积分奖励都由人工设定。这样的公式可读且确定,但它是分段且不可微的,无法随着数据积累而改进,也无法区分真正为零的特征与因管道未能捕获而缺失的特征。我们提出了 zScore-N,即在我们生产环境中取代旧公式的神经网络。该公式充当了它的教师:针对 2019-2024 年间抽样的 5,208,952 个钱包进行校准,并验证其复现生产输出的误差在 2.3e-13 以内,从而在零标签噪声下提供了无限量的带标签训练数据。训练后的网络在 1000 分制上以 0.58 点的均方根误差(R² = 0.99997)复现了该公式,而梯度提升树为 2.25,线性回归在相同特征和划分下为 28.04。使用缺失值掩码对未损坏目标进行训练,它使不完整数据引入的误差减半:在 10% 的特征级缺失率下,公式相对于其自身完整数据输出漂移了 51.4 点,并带有系统性的 -12.5 点偏差,而网络仅漂移 17.9 点。该网络在生产规模上承载评分,覆盖数百万个钱包,其规模和活跃度跨越六个数量级。
英文摘要
Wallet reputation scores decide who receives an airdrop, who can borrow, and who enters an allowlist across decentralised finance. They almost always begin as hand-written formulas: compositions of clamped logarithmic, linear and square-root transforms over behavioural features, with every threshold and point award set by hand. Such a formula is readable and deterministic, but it is piecewise and non-differentiable, it cannot improve as data accumulates, and it cannot distinguish a feature that is genuinely zero from one its pipeline failed to capture. We present zScore-N, the neural network that replaced ours in production. The formula served as its teacher: calibrated against 5,208,952 wallets sampled across 2019-2024 and verified to reproduce production output to within 2.3e-13, it supplies unlimited labelled training data at zero label noise. The trained network reproduces it to 0.58 points RMSE on the 1000-point scale (R^2 = 0.99997), against 2.25 for gradient-boosted trees and 28.04 for linear regression on identical features and splits. Trained with missing-value masks against uncorrupted targets, it halves the error that incomplete data introduces: at 10% feature-level missingness the formula drifts 51.4 points from its own complete-data output with a systematic -12.5 point bias, while the network drifts 17.9. The network carries the score at production scale, across a population of millions of wallets spanning six orders of magnitude in size and activity.