arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

降低跨模型开销:多模型数据上的查询优化

Reducing the Cross-Model Tax: Query Optimization over Multi-Model Data

Jáchym Bártík, Filip Štrobl, Irena Holubová

arXiv 2609.05014首次发表:更新:

发表机构

Charles University(查理大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对跨异构数据模型查询的高开销问题,提出感知映射与能力的优化方法并在 MM-quecat 中实现,可大幅降低查询延迟、消除内存不足故障并缩短规划时间,提升多模型查询处理的效率与鲁棒性。

AI 中文摘要

跨异构数据模型进行查询会因查询分解、数据传输以及底层数据库系统外部处理产生大量开销。本文研究发现,在评估的基于分解的架构中,这种跨模型开销的很大一部分并非源于异构性本身,而是由统一查询处理器做出的可避免决策导致的。本文提出了一种感知映射与能力的优化方法,该方法将处理过程系统性地迁移至靠近数据的位置,结合了感知模型的谓词下推、跨模型依赖连接以及统一优化管道内的非冗余查询部分构建,适用于关系型数据库、文档数据库和图数据库。该方法已在 MM-quecat 中实现,并在 PostgreSQL、MongoDB、Neo4j 及其异构组合上进行了评估,结果显示其可将查询延迟降低最多两个数量级,消除了原始单 DBMS 实验中观察到的所有内存不足故障,通过依赖执行进一步实现了数量级的改进,还将复杂图计划的规划时间从数百毫秒缩短至数毫秒。这些结果表明,已有的优化原则可跨数据模型和系统边界进行推广,能显著提升基于分解的多模型查询处理的效率与鲁棒性。

英文摘要

Querying across heterogeneous data models incurs overhead from query decomposition, result retrieval and conversion, and processing outside the underlying database systems. This paper investigates the extent to which, in a decomposition-based architecture, this cross-model tax results from decisions made by the unifying query processor rather than from heterogeneity alone. We present a mapping- and capability-aware optimization approach that moves applicable processing into native query parts. It combines model-aware predicate pushdown, cross-model dependent joins, and non-redundant query-part construction within a unified pipeline spanning relational, document, and graph databases. The approach is implemented in MM-quecat and evaluated using 20 read-only queries across PostgreSQL, MongoDB, and Neo4j, as well as a heterogeneous combination of the three systems in a single-machine, containerized deployment. For the query--environment combinations most affected by large intermediate results, predicate pushdown yields maximum observed latency reductions of up to two orders of magnitude and prevents the out-of-memory failures observed in the original single-DBMS experiments. Dependent execution further improves eligible external joins, while non-redundant construction reduces planning time for the largest evaluated graph plans, from hundreds of milliseconds to several milliseconds. The results show how established optimization principles can be applied across conceptual, mapping, data-model, and DBMS boundaries in decomposition-based multi-model query processing.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑