arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关系核心图分析:以SQL级别查询图,以及为何节点/边模型是性能负担而非关联数据更真实的图景

Relational-Core Graph Analytics Querying graphs at SQL scale, and why the node/edge model is a performance tax, not a truer picture of connected data

Gene Zhang

arXiv 2609.01525首次发表:更新:

发表机构

ClickGraph; DeltaGraph(ClickGraph; DeltaGraph)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出ClickGraph与DeltaGraph系统,将Cypher转换为原生关系模式在关系引擎执行,性能远超Neo4j等原生图引擎,突破内存图引擎的扩展限制,证明关系系统可高效处理企业图分析工作负载。

AI 中文摘要

长久以来存在一种固有假设:图分析需要专用的图引擎,而关系系统不适合处理关联数据。但我们针对企业实际运行的工作负载提出相反观点:以图查询语言为前端的列式关系引擎,在分析型图查询上可达到或超越原生图引擎的性能,且关键优势在于其扩展性远超内存图引擎失效的临界点。我们进一步指出,节点/边属性图并非关联数据更忠实的模型,而是对关系表中已明确存在的关系的重新编码;在查询时重构这些关系完全是额外开销。我们提出ClickGraph及其Databricks方言的同类系统DeltaGraph,这两个系统可将Cypher直接转换为原生关系模式——即已存在的表、列和外键——并直接在ClickHouse、Databricks或湖仓文件的进程内执行,无需导入数据或单独集群。由于输出是普通SQL,性能不佳的查询有开放的优化空间:可重写查询,也可扩展引擎本身。我们用同行系统已发表的基准测试支撑该论点,其中列式引擎的性能比Neo4j高2至4个数量级,还通过LDBC社交网络基准测试套件的可复现测量结果提供佐证。

英文摘要

A durable assumption holds that graph analytics requires a purpose-built graph engine, and that relational systems are ill-suited to connected data. We argue the opposite for the workloads enterprises actually run. A columnar relational engine fronted by a graph query language matches or exceeds native graph engines on analytical graph queries, and - decisively - scales past the point where in-memory graph engines fail. We further argue that the node/edge property graph is not a more faithful model of connected data but a re-encoding of relationships that already exist explicitly in relational tables; reconstructing them at query time is pure overhead. We present ClickGraph and its Databricks-dialect sibling DeltaGraph, systems that translate Cypher directly onto the native relational schema - the tables, columns, and foreign keys as they already exist - and execute in place on ClickHouse, Databricks, or in-process on lakehouse files, with no import and no separate cluster. Because the output is ordinary SQL, an underperforming query is an open optimization surface: it can be rewritten, and the engine itself extended. We support the argument with a peer system's own published benchmark, in which a columnar engine outruns Neo4j by two-to-four orders of magnitude, and with reproducible measurements across the LDBC Social Network Benchmark suite.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑