AI 中文总结
本文针对天际线查询优化问题,提出VoS变量排序策略,通过优化属性排序减少冗余比较,经多类数据集与架构实验验证可降低计算开销、提升查询性能。
AI 中文摘要
天际线算法的效率高度依赖底层数据特征。传统优化工作聚焦于最小化元组对支配性检查的总数以提升查询性能,但实践中,两个元组间的支配性检查未必需要评估数据的每个偏好属性的支配关系,这导致支配性检查优化与查询执行性能存在脱节。本文提出,天际线算法需同时优化每个元组的支配性检查总数和每个属性的支配性检查总数,且属性(或变量)的排序对这两个目标的天际线计算效率有显著影响。基于该前提,本文提出多种策略以确定有效变量排序,从而减少冗余属性比较。在合成与真实世界数据集、标量与SIMD架构上开展的大量实验证实,所提方法可降低计算开销、提升天际线查询性能。
英文摘要
Efficiency of skyline algorithms is highly influenced by the underlying data characteristics. Traditionally, optimization efforts have focused on minimizing the total number of tuple-pair dominance checks to improve query performance. However, in practice, a dominance check between two tuples does not necessarily require evaluating dominance relationships for each and every preference attribute of the data and this creates a disconnect between dominance checks optimization and query execution performance. In this paper, we argue that skyline algorithms need to optimize total per-attribute dominance checks, along with per-tuple dominance checks and that, for both of these goals, the ordering of the attributes (or variates) can have a substantial impact on the efficiency of skyline computation. Based on this premise, we present several strategies for identifying an effective variate order to minimize redundant attribute comparisons. Extensive experiments on both synthetic and real-world datasets, and on both scalar and SIMD architectures, confirm the effectiveness of the proposed approach in reducing computational overhead and improving skyline query performance.