发表机构
Amazon; Rutgers University(亚马逊; 罗格斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究多语言编码器是否能为不同语言的产品生成一致的语义ID,发现多语言暴露和平衡码本使用不足以保证跨语言SID一致性。
AI 中文摘要
语义ID(SIDs)将项目嵌入压缩为用于生成式检索的离散代码序列。我们探究多语言编码器是否足以使同一产品的不同语言渲染获得语言一致的语义ID。使用以英语、西班牙语和日语渲染的Amazon ESCI列表,我们测试翻译是否保持与英语原文接近,残差量化是否对翻译引起的移动异常敏感,以及多语言或语言平衡的量化器拟合是否能提高SID一致性。多语言E5使翻译明显分离:在英语为主的拟合下,日语翻译仅7.7%的情况保留其英语对应物的第一个SID代码,而英语改写为89.0%。距离匹配的产品导向对照产生与翻译几乎相同的完整SID不匹配,没有证据表明量化器选择性地放大语言方向。平衡拟合混合使码本使用更均匀,但进一步降低跨语言前缀一致性:西班牙语第一代码一致性从28.3%降至6.6%,而仅英语拟合为67.6%的西班牙语翻译保留该一致性。这些结果表明,仅多语言暴露和平衡码本使用并不能保证语言一致的SID。
英文摘要
Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spanish, and Japanese, we test whether translations remain close to their English source, whether residual quantization is unusually sensitive to translation-induced movement, and whether multilingual or language-balanced quantizer fitting improves SID agreement. Multilingual E5 places translations measurably apart: under an English-heavy fit, a Japanese translation preserves the first SID code of its English counterpart in only 7.7% of cases, compared with 89.0% for an English rewording. Distance-matched product-directed controls produce nearly the same full-SID mismatch as translation, providing no evidence that the quantizer selectively amplifies language directions. Balancing the fitting mixture makes codebook use more uniform but further reduces cross-lingual prefix agreement: Spanish first-code consistency falls from 28.3% to 6.6%, while an English-only fit preserves it for 67.6% of Spanish translations. These results show that multilingual exposure and balanced codebook use alone do not guarantee language-consistent SIDs.
Comments7 pages, 8 tables. Accepted as a short paper at WiNLP 2026, co-located with EMNLP 2026