发表机构
University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出锚点散度方法,通过对比学习、指数族与信息几何的对应,在固定表示上定义上下文相关的语义几何,实验证明其能高效指定特定语义相似性。
AI 中文摘要
本文关注语义上下文如何决定学习到的向量表示中的几何结构。相似性通常使用余弦相似度来衡量,它提供了一种固定的几何结构。然而,语义相似性本质上是上下文相关的:两幅图像可能因为描绘同一物体、共享视觉风格或与同一临床发现相关而相似。我们表明,对比表示自然包含一族几何结构,这些几何结构可以专门针对特定的语义结构。关键思想是利用对比学习、指数族和信息几何之间的相互作用,在“锚点”上的概率分布与表示空间上的Bregman几何之间建立对应关系。我们利用这种对应关系定义了“锚点散度”,一种在固定表示上指定特定于上下文的语义几何的方法。在这种对应关系下,对锚点分布进行建模即对几何本身进行建模。检索实验表明,锚点散度为指定特定于上下文的语义相似性提供了一种有效且高效的方法。
英文摘要
This paper concerns how semantic context determines geometry in learned vector representations. Similarity is typically measured using cosine similarity, which provides a single fixed geometry. Semantic similarity, however, is inherently context dependent: two images may be similar because they depict the same object, share a visual style, or are relevant to the same clinical finding. We show that contrastive representations naturally encompass a family of geometries that can be specialized to particular semantic structure. The key idea is to use an interplay between contrastive learning, exponential families, and information geometry to establish a correspondence between probability distributions over "anchors" and Bregman geometries on the representation space. We use this correspondence to define "Anchor Divergences", a method for specifying context-specific semantic geometries on fixed representations. Under this correspondence, modeling the anchor distribution models the geometry itself. Experiments on retrieval show that anchor divergences provide an effective and efficient way to specify context-specific semantic similarity.
CommentsCode is available at https://github.com/sky1712/Anchor-Divergence