arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11035cs.DB

流数据上的连续查询:用于Top-$K$最大和区间

Continuous Query for Top-$K$ Maximal Sum Intervals over Streaming Data

Zhongshuai Zhang, Xiaochun Yang, Baihua Zheng, Rui Zhu, Haomin Li, Bin Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对流数据上连续查询Top-$K$最大和区间的问题,提出基于分区的策略,通过特定分区方案确保最大和区间在单个分区内,实现独立并行处理,开发相关算法,经实验验证显著提升了效率。

中文摘要 AI 辅助

使用滑动窗口在数据流上连续识别Top-$k$最大和区间,对物联网及其他应用至关重要。最大和区间是有符号值序列中具有最大和的非重叠连续子序列。现有算法不适用于流环境:要么即使对于小$k$值也会详尽枚举所有区间,要么依赖需要频繁且昂贵重构的索引。我们提出一种基于分区的新策略。核心见解是一种分区方案,确保任何最大和区间完全包含在单个分区内,实现独立和并行处理。该设计有两个关键优势:能安全修剪对Top-$k$结果无贡献的分区,大幅缩小搜索空间;能高效、增量维护每个分区中的最大和区间。我们开发了分区构建、增量分区更新和基于分区的Top-$k$最大和区间搜索算法。在真实和合成数据集上的大量实验表明,我们的方法显著提高了效率。

英文摘要

The continuous identification of top-$k$ maximal sum intervals using a sliding window over a data stream is a critical operation for applications in IoT and beyond. A maximal sum interval is a non-overlapping, contiguous subsequence with the maximal sum in a sequence of signed values. Existing algorithms are ill-suited for streaming contexts: they either exhaustively enumerate all intervals even for small $k$ values, or depend on indexes that require frequent and costly restructuring. We propose a novel partition-based strategy. Our core insight is a partitioning scheme that guarantees that any maximal sum interval is fully contained within a single partition, enabling independent and parallel processing. This design provides two key advantages: it enables safe pruning of partitions that cannot contribute to top-$k$ results, drastically narrowing the search space, and it enables efficient, incremental maintenance of the maximal sum intervals in each partition. We develop algorithms for partition construction, incremental partition updates, and partition-based top-$k$ maximal sum interval search. Extensive experiments on real and synthetic datasets demonstrate that our approach significantly improves efficiency.

补充信息

↑