AI 中文总结
本文针对PostgreSQL-V 1.0的三个局限推出PostgreSQL-V 2.0,实现并发向量搜索更新、快速崩溃恢复与物理复制,吞吐量达前者36.4倍,恢复时间约20毫秒,是内嵌PostgreSQL的集成向量数据库系统。
AI 中文摘要
本文提出了PostgreSQL-V 2.0,这是一种内嵌于PostgreSQL的可扩展集成向量数据库系统。现有的基于PostgreSQL的向量搜索系统如pgvector,将向量索引嵌入PostgreSQL面向页面的存储引擎,产生了显著的开销,导致与专用向量数据库存在巨大的性能差距。在我们早期的工作中,我们推出了PostgreSQL-V 1.0,它通过将向量索引结构与PostgreSQL的存储引擎分离来解决该问题,在保持SQL兼容性的同时,使向量搜索性能接近原生向量索引库。然而,我们发现PostgreSQL-V 1.0存在三个对实际工作负载至关重要的局限:仅支持单个连接(无并发)、恢复时间随索引大小增长、不支持物理复制。我们进一步提出了PostgreSQL-V 2.0,它弥补了这三个不足。PostgreSQL-V 2.0的并发支持实现了PostgreSQL多进程后端间完全并发的向量搜索和更新,在服务32个并发客户端时,吞吐量达到PostgreSQL-V 1.0的36.4倍。PostgreSQL-V 2.0的快速崩溃恢复使成本与总索引大小无关,维持在约20毫秒,而PostgreSQL-V 1.0的恢复时间增长至秒级。PostgreSQL-V 2.0的物理复制支持将物理复制扩展至解耦的索引,在不增加主节点负担的前提下保持备节点的索引一致性。综上,这些进展使PostgreSQL-V 2.0成为内嵌于PostgreSQL的完全并发、抗崩溃且支持复制的向量数据库。
英文摘要
This paper presents PostgreSQL-V 2.0, a scalable integrated vector database system inside PostgreSQL. Existing PostgreSQL-based vector search systems such as pgvector embed vector indexes into PostgreSQL's page-oriented storage engine, incurring significant overhead that leads to a huge performance gap with specialized vector databases. In our earlier work, we introduced PostgreSQL-V 1.0, which addresses this issue by separating vector index structures from PostgreSQL's storage engine, enabling vector search performance close to that of native vector index libraries while preserving SQL compatibility. However, we find that PostgreSQL-V 1.0 has three limitations that matter for real-world workloads: it only supports a single connection (without concurrency), recovery time grows with index size, and physical replication is unsupported. We further present PostgreSQL-V 2.0, which closes all three gaps. PostgreSQL-V 2.0's concurrency support enables fully concurrent vector searches and updates across PostgreSQL's multi-process backends, delivering up to 36.4x the throughput of PostgreSQL-V 1.0 while serving 32 concurrent clients. PostgreSQL-V 2.0's fast crash recovery keeps cost independent of total index size, remaining near 20 ms while PostgreSQL-V 1.0's grows into seconds-scale. PostgreSQL-V 2.0's physical replication support extends physical replication to the decoupled index, preserving index consistency on standbys without burdening the primary node. Together, these advances make PostgreSQL-V 2.0 a fully concurrent, crash-resilient, and replication-ready vector database inside PostgreSQL.