发表机构
University of Zürich(苏黎世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将句子对的词元嵌入视为Transformer潜空间的点云,探究Dowker同调捕捉句子相似度的能力,发现其可用于相似度评估与可视化,推导的单值摘要表现尚可但未优于标准度量。
AI 中文摘要
Dowker同调是一种拓扑工具,可用于分析处于同一空间中的两个点云的相对位置。本研究将构成句子对的词元嵌入视为Transformer模型潜空间中的一对点云,分别使用针对句子相似度微调过和未微调过的模型,探究Dowker同调是否能捕捉句子相似度信息。我们通过将Dowker同调特征回归到真实相似度分数上进行评估,发现Dowker同调确实能捕捉句子相似度信息,且可用于相似度数据与模型的可视化检查。为使Dowker同调便于应用,我们从中推导了单值摘要,期望其能直接捕捉句子相似度,这些单值摘要表现尚可,但未优于基于成熟池化方法的标准句子相似度度量。
英文摘要
Dowker homology is a topological tool that may be used to analyze the relative position of two point clouds living in a common space. We investigate whether Dowker homology captures sentence similarity information by treating the embeddings of the tokens that constitute a sentence pair as a pair of point clouds in the latent space of a transformer model, using both models that have and have not been fine-tuned for sentence similarity. We find that Dowker homology captures sentence similarity information, as measured by regressing Dowker homology features onto ground-truth similarity scores, and that it can be used for visual inspection of similarity data and models. In an attempt to make Dowker homology readily applicable, we derive from it single-number summaries that we expect to capture sentence similarity directly. These turn out to work reasonably well, but without outperforming standard sentence similarity measures based on established pooling methods.
Comments12 pages; accepted at "Workshop on Insights from Negative Results in NLP 2026"