arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文轨迹与增量上下文位移:利用LLM理解动态的、话语特定的意义构建

Contextual trajectory and incremental contextual displacement: Towards using LLMs to understand dynamic, utterance-specific meaning construction

Grayson Wycliffe Storer, Julia Witte Zimmerman

arXiv 2610.00840首次发表:更新:

发表机构

University of Vermont; Vermont Complex Systems Institute(佛蒙特大学; 佛蒙特复杂系统研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出逐标记增量轨迹方法,通过反复计算上下文词嵌入来追踪话语展开过程,在花园路径句上验证其能区分歧义句并揭示话语级信息分布,为利用LLM研究动态意义构建提供了新框架。

AI 中文摘要

基于Transformer的大型语言模型(LLM),如RoBERTa,使用上下文词嵌入(CWE)来表示文本,这些嵌入会根据周围上下文改变每个标记的嵌入。我们通过反复重新计算一个标记的CWE,随着句子中连续添加单词,构建逐标记的增量轨迹,从而表示上下文嵌入在话语展开过程中如何演变。我们使用花园路径句作为具有特征性测试案例来评估这种方法。逐标记轨迹再现了花园路径处理的已知特征,包括关键区域周围的干扰,并能可靠地将花园路径句与匹配的消歧控制句区分开来。我们引入了几个量化跨上下文增量表示位移的指标,并表明轨迹信息可以高度预测句子类型。我们发现,与歧义相关的信息不仅可以从句子级别的CLS表示中恢复,还可以从普通词汇标记中恢复,这表明话语级别的信息分布在多个表示尺度上。在探索性分析中,我们在其他与歧义和误导相关的语言现象中发现了定性相似的轨迹结构。总之,这些结果确立了逐标记增量轨迹作为利用LLM研究话语特定意义构建的一个有前景的框架。

英文摘要

Transformer-based large language models (LLMs) such as RoBERTa represent text using contextual word embeddings (CWEs), which alter the embeddings associated with each token based on surrounding context. We construct token-wise incremental trajectories by repeatedly recomputing a token's CWE as successive words are added to a sentence, yielding a representation of how contextualized embeddings evolve as the utterance unfolds. We evaluate this approach using garden-path sentences as a test case with characteristic features. Token-wise trajectories reproduce known features of garden-path processing, including disruption around the critical region, and reliably distinguish garden-path sentences from matched disambiguated controls. We introduce several metrics for quantifying representational displacement across contextual increments and show that trajectory information can be highly predictive of sentence type. We find that ambiguity-related information is recoverable not only from the sentence-level CLS representation but also from ordinary vocabulary tokens, suggesting that utterance-level information is distributed across multiple representational scales. In exploratory analyses, we find qualitatively similar trajectory structures in other ambiguity- and misdirection-related linguistic phenomena. Together, these results establish token-wise incremental trajectories as a promising framework for studying utterance-specific meaning construction using LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑