用于动态时间感知最近邻搜索的版本化统一图索引
A Versioned Unified Graph Index for Dynamic Timestamp-Aware Nearest Neighbor Search
中文总结 AI 辅助
研究针对动态向量数据集的时间感知最近邻搜索问题,提出TiGER索引方法,通过版本化连通性构建统一图实现高效查询,QPS最高提升5倍且不损失准确率,可用于实时推荐等场景。
中文摘要 AI 辅助
我们提出了TiGER(Time-Integrated Graph for Efficient Retrieval,高效检索的时间集成图),这是一种针对动态向量数据集执行快速时间感知近似最近邻搜索的新方法,支持对任意可能的时间范围灵活操作。所提算法通过利用基于集成版本化连通性的索引结构,为所有向量构建并维护统一图,允许在该统一图上直接查询任意时间间隔,无需遍历无效向量,从而省去了搜索后过滤、合并或为每个可能的复合范围单独构建图的需求。实证评估表明,我们的方法在不降低基于过滤或按时间片段子图的基准方法准确率的前提下,查询每秒处理量(QPS)最高提升了5倍。我们认为该方法将支持实时推荐系统、日志分析及任何需要对动态、时间分段数据进行快速相似性搜索的场景中,对不断演变的数据集开展高效时间分析。
英文摘要
We present TiGER (Time-Integrated Graph for Efficient Retrieval), a novel approach for performing fast time-aware approximate nearest neighbor searches on dynamic vector datasets with flexibility over any possible time range. Our proposed algorithm builds and maintains a unified graph for all vectors by leveraging an index structure based on integrated versioned connectivity, allowing arbitrary time intervals to be queried directly on the unified graph without having to traverse invalid vectors. This forgoes the need for post-search filtering or merging, or separate graphs for each possible composite range. Empirical evaluations show that our method attains up to a 5x improvement in queries per second (QPS) without compromising accuracy over baselines based on filtering or per-time-segment sub-graphs. We believe that this method will enable efficient temporal analysis across evolving datasets in real-time recommendation systems, log analysis, and any scenario requiring fast similarity search over dynamic, time-segmented data.