arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于高效稠密向量检索的超图嵌入索引

Hypergraph Embedding Indexing for Efficient Dense Vector Retrieval

Kishore Konda

arXiv 2608.22980首次发表:更新:

发表机构

Sodhana(索达纳)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有ANN索引将嵌入视为高维不可分割点的问题,提出超图嵌入索引(HEI)框架,通过多互补超图提升检索覆盖率,引入激活多样性作为诊断指标。

AI 中文摘要

稠密向量检索已成为现代语义搜索的基础,但现有的近似最近邻(ANN)索引将嵌入视为高维空间中不可分割的点。本研究提出超图嵌入索引(HEI)这一框架,转而依据高度激活的潜在嵌入维度的组合来组织文档,该公式支持倒排索引式的候选生成,同时保留稠密嵌入的语义排序能力。我们进一步证明,构建多个互补超图可大幅提升检索覆盖率,且不会出现单个超图维度增加带来的组合式增长问题。最后,我们证实嵌入激活的统计特性强烈影响坐标倒排索引的效率,引入“激活多样性”作为坐标倒排框架中控制嵌入可索引性的诊断指标。

英文摘要

Dense vector retrieval has become the foundation of modern semantic search, yet existing approximate nearest neighbor (ANN) indexes treat an embedding as an indivisible point in a high-dimensional space. In this work, we propose the Hypergraph Embedding Index (HEI), a framework that instead organizes documents according to combinations of highly activated latent embedding dimensions. This formulation enables inverted-index style candidate generation while preserving the semantic ranking capabilities of dense embeddings. We further demonstrate that constructing multiple complementary hypergraphs substantially improves retrieval coverage without the combinatorial growth associated with increasing the dimensionality of a single hypergraph. Finally, we establish that the statistical properties of embedding activations strongly influence coordinate-inverted indexing efficiency, introducing \emph{activation diversity} as a diagnostic metric governing embedding indexability in coordinate-inverted frameworks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑