AI 中文总结
RVANNS是面向RISC-V向量扩展的ANNS引擎,通过混合精度多层索引与ROrder优化,在RVV处理器及Milvus中实现了ANNS的显著加速,性能优于相关CPU与GPU基准。
AI 中文摘要
CPU上的近似最近邻搜索(ANNS)越来越多地受限于候选向量移动与解码,而非峰值算术吞吐量。尽管RISC-V向量扩展(RVV)提供了向量长度无关的执行方式以及基于LMUL的寄存器分组,但通用低精度解码仍会产生转换开销,而不规则图遍历会生成分散的访问操作,降低缓存局部性与内存级并行性。我们提出RVANNS,一种面向RVV的ANNS引擎,联合优化向量表示与图局部性。其混合精度多层索引(MPMI)用密集8位仿射基和稀疏FP16/FP32残差表示每个向量,将重构与距离累积融合,并将扩展对齐LMUL大小的寄存器组。ROrder将可能共同访问的图节点共置,对重映射的邻接列表排序,将分散的有效载荷探测转化为更密集、主要向前移动的地址流。集成到Milvus后,RVANNS在实际128位和256位RVV处理器上分别实现了比标量执行高3.39倍和4.94倍的加速。在受控HNSW配置下,它的吞吐量比RVV SIMD+FP32高2.27至2.76倍,比对应的AVX-512和SVE基准高1.18至1.59倍。在Cohere10M数据集上,它的每瓦QPS比所评估的GPU基准高1.82至2.27倍。
英文摘要
Approximate nearest neighbor search (ANNS) on CPUs is increasingly constrained by candidate-vector movement and decoding rather than peak arithmetic throughput. Although the RISC-V Vector Extension (RVV) provides vector-length-agnostic execution and LMUL-based register grouping, generic low-precision decoding still incurs conversion overhead, while irregular graph traversal generates scattered accesses that degrade cache locality and memory-level parallelism. We present RVANNS, an RVV-oriented ANNS engine that jointly optimizes vector representation and graph locality. Its Mixed-Precision Multi-Layer Index (MPMI) represents each vector with a dense 8-bit affine base and sparse FP16/FP32 residuals, fusing reconstruction with distance accumulation and aligning widening with LMUL-sized register groups. ROrder co-locates likely co-visited graph nodes and sorts remapped adjacency lists, transforming scattered payload probes into denser, predominantly forward-moving address streams. Integrated into Milvus, RVANNS achieves 3.39x and 4.94x speedups over scalar execution on real 128-bit and 256-bit RVV processors, respectively. Under controlled HNSW configurations, it improves throughput by 2.27--2.76x over RVV SIMD+FP32 and by 1.18--1.59x over the corresponding AVX-512 and SVE baselines. On Cohere10M, it further delivers 1.82--2.27x higher QPS/W than the evaluated GPU baselines.