发表机构
University of Tübingen; University of Amsterdam(图宾根大学; 阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出SAGE框架,结合认知模型与语言模型,将语用过程分解为提议者、评估者和选择器模块。通过三个案例研究评估,该模型取得高精度且常超基线,不过组件分析有不对称性,还探讨了神经符号模型在语用语言使用中的前景与局限。
AI 中文摘要
语用语言使用需要对替代方案进行推理:说话者可能选择的替代表达,或听者可能考虑的替代解释。因此,语用学的形式和计算模型必须指定对话者推理的替代方案集,通常通过手动指定来完成。在此,我们提出了一个框架,即用于解释的脚手架生成模型(SAGE),它将认知模型的解释透明度与语言模型(LMs)的生成灵活性相结合。SAGE将语用过程分解为三种模块:提议者,使用LMs生成候选替代方案的开放式空间;评估者,评估这些替代方案(例如它们的语义、复杂性或典型性);选择器,执行基于规则的认知动机任务分析的计算步骤。我们在三个案例研究中评估SAGE,涵盖语用生成和解释——指代表达生成、方式(M-)含义和格莱斯会话含义。使用计算认知建模的既定方法对SAGE模型进行严格评估,包括消融、基线比较和与人类数据的定量拟合。在各项研究中,SAGE模型取得了高精度,且常常优于基线,但组件级分析揭示了一种不对称性:LM提议者可靠地生成了适合语用建模的替代方案,而LM评估者更擅长提供直观判断而非理论或形式度量的判断。我们讨论了神经符号模型作为人类语用语言使用的候选解释性说明的前景和局限性。
英文摘要
Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must therefore specify the sets of alternatives that interlocutors reason over, which is often done through manual specification. Here we propose a framework, ScAffolded Generative models for Explanation (SAGE), that combines the explanatory transparency of cognitive models with the generative flexibility of language models (LMs). SAGE decomposes a pragmatic process into three kinds of modules: proposers, which use LMs to generate an open-ended space of candidate alternatives; evaluators, which assess those alternatives (e.g., their semantics, complexity, or typicality); and selectors, which implement the rule-based computational steps of a cognitively motivated task analysis. We assess SAGE in three case studies spanning pragmatic generation and interpretation-referential expression generation, manner (M-)implicatures, and Gricean conversational implicatures. SAGE models are evaluated critically using established methods from computational cognitive modeling, including ablations, baseline comparisons, and quantitative fit to human data. Across studies, SAGE models achieved high accuracy and often outperformed baselines, but component-level analyses reveal an asymmetry: LM proposers reliably generated alternatives well-suited to pragmatic modeling, whereas LM evaluators are better at providing intuitive judgements rather than judgements of theoretical or formal measures. We discuss the promise and the limitations of neuro-symbolic models as candidate explanatory accounts of human pragmatic language use.
Comments27 pages main text, 9 figures; 25 pages supplementary materials