发表机构
University of Lausanne(洛桑大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出符号学感知框架,通过保真度和覆盖率两指标评估文本间符号学对齐,发现覆盖率低于保真度且低温下LLM与人工数据对齐最佳。
AI 中文摘要
当两段文本描述同一表达时,基于词汇重叠或整篇文本相似度的标准指标可能无法检测出该表达在框架方式上的有意义差异。我们提出一个评估文本间符号学对齐的框架,其中符号学轮廓涵盖文本所凸显的语境意义和话语指涉。我们的方法产生两个分数,即符号学保真度和符号学覆盖率,分别估计一段文本的轮廓在多大程度上被另一段文本支持,以及它在多大程度上恢复了另一段文本的轮廓。实验表明,覆盖率通常低于保真度,并且大型语言模型与人工策展数据之间的对齐在低采样温度下最高,而较高温度会降低这种对齐。
英文摘要
When two texts describe the same expression, standard metrics based on lexical overlap or whole-text similarity may fail to detect meaningful differences in how that expression is framed. We propose a framework to evaluate semiotic alignment between texts, where a semiotic profile encompasses both the contextual meaning and the discourse references made salient by a text. Our approach yields two scores, Semiotic Fidelity and Semiotic Coverage, estimating how much of one text's profile is supported by the other and how much of the other's profile it recovers. Experiments show that coverage is typically lower than fidelity, and that alignment between LLMs and human-curated data is highest at low sampling temperatures, while higher temperatures reduce this alignment.