arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探索自底向上聚类以创建语义ID

Exploring Bottom-Up Clustering for Creating Semantic IDs

Leah Woldemariam, Sudhanshu Garg, Taha Belkhouja, Charles Kim-Yip, Ali Sahami

arXiv 2609.08310首次发表:更新:

发表机构

Cornell University; PayPal(康奈尔大学; 贝宝)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种自底向上聚类算法生成语义ID,确保标识符唯一且保留嵌入结构,提升聚类质量与下游生成式检索性能。

AI 中文摘要

生成式检索的成功在很大程度上归因于语义ID的使用,语义ID通过捕捉物品的语义,优于哈希等任意物品级标识符。然而,构建语义ID时面临的主要挑战在于将每个标识符映射到唯一的产品,并捕获对下游任务有价值的信息。以往的工作通过附加额外的码字来去重物品标识符,并利用残差量化来创建层次聚类。在本工作中,我们提出了一种生成语义ID的算法,确保标识符既唯一又保留原始嵌入的结构。我们工作的关键在于使用自底向上聚类来保留嵌入空间中的局部结构,从而提高所得语义ID的聚类质量及其在下游生成式检索中的实用性。

英文摘要

The success of generative retrieval has largely been attributed to the use of Semantic IDs, which improve over arbitrary item-level identifiers such as hashes by capturing the semantics of items. The main challenges faced when constructing Semantic IDs, however, is in mapping each identifier to a unique product and capturing information valuable to downstream tasks. Past works have appended additional codewords to de-duplicate item identifiers and utilized residual quantization to create hierarchical clusters. In this work, we present an algorithm for generating Semantic IDs that ensure the identifiers are both unique and preserve the structure of the original embedding. Key to our work is the use of bottom-up clustering to preserve local structure in the embedding space, improving the clustering quality of the resulting Semantic IDs and their utility for downstream generative retrieval.

Comments6 Pages, workshop Paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑