arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29173cs.DB

MERIT:基于动态图的近似最近邻索引的高效原位删除方法

MERIT: Efficient In-Place Deletion for Dynamic Graph-Based Approximate Nearest Neighbor Indexes

Zekai Wu, Jiabao Jin, Peng Cheng, Wangze Ni, Haoyang Li, Lei Chen, Junjie Yao, Jingkuan Song, Heng Tao Shen

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对动态图近似最近邻索引的原位删除难题,提出MERIT框架,通过三项核心技术实现高效删除,集成于两种主流索引,实验显示其删除速度远优于SOTA方法且召回率稳定。

中文摘要 AI 辅助

图索引已成为高维数据近似最近邻搜索(ANNS)的主流方法,在检索增强生成、推荐系统、向量数据库等实际应用中发挥关键作用。尽管静态图构建与搜索已取得大量进展,但高效原位删除仍具挑战性:需移除过时向量,同时避免失效入边占用搜索资源,或因全局图维护中断检索增强生成(RAG)、推荐平台等在线服务。为解决该问题,本文提出MERIT(基于最小生成树的高效原位更新修复方法),这一原位更新框架包含三项核心技术:(1)基于有界搜索的恢复,结合被删除顶点的出邻接点与易搜索的入邻接点;(2)kr-最小生成树(MST)局部修复,在提升局部连通性的同时为图搜索保留多条路由选择;(3)带版本的边失效机制,立即过滤指向被删除顶点的所有失效入边,并在重写邻接表时逐步移除它们。将其与分层HNSW索引、单层Vamana索引集成,证明其适用于不同图结构。在多个真实数据集上的大量实验显示,MERIT处理删除的成本接近插入1个向量的成本,比现有最优(SOTA)方法快3.02倍至18.87倍,且随着删除操作累积,搜索召回率保持稳定甚至提升。

英文摘要

Graph-based indexes have become the dominant approach to approximate nearest neighbor search (ANNS) over high-dimensional data and play a crucial role in real-world applications such as retrieval-augmented generation, recommendation systems, and vector databases. Despite extensive progress in static graph construction and search, efficient in-place deletion remains challenging because obsolete vectors must be removed without allowing stale incoming edges to consume search capacity or expensive graph-wide maintenance to interrupt online services, e.g., retrieval-augmented generation (RAG) and recommendation platforms. To address this problem, we propose MERIT (MST-based Efficient Repair with In-place updaTes), an in-place update framework with three core techniques: (1) bounded search-based recovery that combines a deleted vertex's outgoing neighbors with its readily searchable in-neighbors, (2) $k_r$-Minimum Spanning Tree (MST) local repair that promotes local connectivity while retaining multiple routing choices for graph search, and (3) versioned-edge invalidation that immediately filters all stale incoming edges to the deleted vertex and progressively removes them as adjacency lists are rewritten. Its integration with the hierarchical HNSW index and the single-layer Vamana index demonstrates applicability across distinct graph structures. Extensive experiments on multiple real-world datasets show that MERIT processes deletion at nearly the cost of inserting one vector, achieves up to $3.02\times$--$18.87\times$ faster deletion than state-of-the-art (SOTA) methods, and keeps search recall stable or even improves it as deletions accumulate.

补充信息

↑