AI 中文总结
该研究针对流处理引擎滑动窗口聚合的归因问题,提出了谓词级夏普利值的闭式计算方法,实现了高效精确的谓词级归因,且在大规模数据集上验证了其速度与精度优势。
AI 中文摘要
流处理引擎会实时报告滑动窗口聚合结果,但无法解释聚合值为何取当前值。一个自然的解决方案是采用合作博弈论中的夏普利值,它可从公理层面将聚合结果分配给窗口内的元组。然而,从业者提出的是谓词级问题(例如,某一区域或客户层级对平均值或方差峰值的贡献有多大)。精确夏普利计算的复杂度随窗口大小呈指数增长,且现有估计器会忽略连续窗口间的大量重叠。我们证明,对于 SUM(求和)、COUNT(计数)、AVG(平均值)以及总体/样本方差,精确谓词级夏普利值可通过每个谓词的三个可加维护摘要(计数、求和、平方和)以闭式形式表达,其系数仅取决于两个运行调和数。因此,归因可简化为每个滑动步长对已注册谓词进行 O(1) 级别的摘要更新,无需枚举联盟。重叠谓词和复合谓词可通过布尔签名的原子细化精确求解。我们进一步表征该现象:每个矩多项式聚合都具有此类闭式形式,而 MAX(最大值)、MIN(最小值)和分位数在任何固定矩阶下均不存在此类形式。实验在超过 10,000 个窗口上,将夏普利值与暴力计算结果匹配至浮点精度,在 N=10^6 时每个滑动步长维持约 2 微秒的处理速度(比相同公式的逐窗口重新计算快 3,200 倍),并以每秒约 180 万次摘要更新的速度解释了纽约市 290 万次出租车行程的夜间车费峰值。
英文摘要
Streaming engines report sliding-window aggregates in real time, but they do not explain \emph{why} an aggregate takes its current value. A natural target is the Shapley value from cooperative game theory, which axiomatically distributes an aggregate among the tuples in the window. Practitioners, however, ask predicate-level questions (e.g., how much a region or customer tier contributed to an average or variance spike). Exact Shapley computation is exponential in the window size, and existing estimators discard the massive overlap between consecutive windows. We show that for SUM, COUNT, AVG, and population/sample variance, exact predicate-level Shapley values admit closed forms in three additively maintained summaries per predicate (count, sum, and sum of squares), with coefficients that depend only on two running harmonic numbers. Attribution therefore reduces to $O(1)$ summary updates per slide for registered predicates, with no coalition enumeration. Overlapping and compositional predicates are answered exactly via atomic refinement of Boolean signatures. We further characterize the phenomenon: every moment-polynomial aggregate admits such a form, while MAX, MIN, and quantiles provably do not at any fixed moment order. Experiments match brute-force Shapley values to floating-point precision on over $10{,}000$ windows, sustain $\approx\!2\,μ$s per slide up to $N=10^6$ ($3{,}200\times$ faster than per-window recomputation of the same formulas), and explain a nighttime fare spike on 2.9M NYC taxi trips at $\approx\!1.8$M summary updates per second.
Commentssubmitted to IEEE Access