arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26930cs.DB

增量Delta-Shapley:用于滑动窗口上谓词归因的独立运行时

Incremental Delta-Shapley: A Standalone Runtime for Predicate Attribution on Sliding Windows

Pouya Khani, Ira Assent

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出独立运行时IDS,解决滑动窗口连续聚合查询的谓词归因问题,其增量维护高效、归因精度高,能显著降低临时计算成本。

中文摘要 AI 辅助

针对滑动窗口的连续聚合查询在实时分析中十分常见,但大多数系统仅报告聚合结果“是什么”,而不说明“哪些谓词”对结果起作用。一篇配套论文[khani2026closedformpredicatelevelshapleyattribution]表明,SUM、COUNT、AVG和方差的精确谓词级Shapley归因仅需三个带闭式系数的加性谓词摘要。这些结果解决了数学问题,但未涉及运行时如何在窗口滑动间维护摘要、暴露归因、响应未注册谓词或摊销重复临时计算的问题。我们提出了IDS(Incremental Delta-Shapley,增量Delta-Shapley),这是一个独立的单节点运行时,可将这些闭式公式转化为可部署的解释系统。IDS消耗窗口维护增量,更新全局、边际和原子摘要,并以常数时间评估任何闭式公式;重叠谓词使用原子细化,受限类SQL API将归因及其逐滑动变化作为一等算子暴露。未注册谓词可通过保留状态扫描、倒排索引或具有浓度保证的摊销滑动窗口采样来响应;频繁谓词可通过重建细化来升级。在合成、对抗、NEXMark风格和纽约出租车工作负载上,归因与穷举Shapley枚举匹配至浮点精度;增量维护的时间复杂度与N成线性关系,比同形式的逐窗口扫描快达4.3×10^5倍;自适应升级在Zipfian轨迹上将临时成本降低达9.2倍。

英文摘要

Continuous aggregate queries over sliding windows are common in real-time analytics, but most systems report \emph{what} an aggregate is doing without attributing \emph{which} predicates account for the result. A companion paper~\cite{khani2026closedformpredicatelevelshapleyattribution} shows that exact predicate-level Shapley attribution for SUM, COUNT, AVG, and variance needs only three additive predicate summaries with closed-form coefficients. Those results settle the mathematics, not how a runtime maintains summaries across slides, exposes attribution, answers unregistered predicates, or amortizes repeated ad hoc ones. We present \textbf{IDS} (Incremental Delta-Shapley), a standalone single-node runtime that turns those closed forms into a deployable explanation system. IDS consumes window-maintenance deltas, updates global, marginal, and atom summaries, and evaluates any closed form in constant time. Overlapping predicates use atomic refinement, and a restricted SQL-like API exposes attribution and its per-slide change as first-class operators. Unregistered predicates are answered by a retained-state scan, an inverted index, or an amortized sliding-window sample with concentration guarantees; frequent ones are promoted by rebuilding the refinement. On synthetic, adversarial, NEXMark-style, and NYC taxi workloads, attribution matches exhaustive Shapley enumeration to floating-point precision; incremental maintenance is flat in $N$ and up to $4.3\times10^{5}\times$ faster than per-window scans of the same form; and adaptive promotion cuts ad hoc cost by up to $9.2\times$ on Zipfian traces.

↑