arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22180q-bio.NC

话语循环、言语不流畅与主动推理

Cycles of Discourse, Speech Dysfluency, and Active Inference

Thomas Parr, Birtan Demirel, Youssuf Saleh, Sanjay Manohar

首次发表
浏览论文内容

中文总结 AI 辅助

研究言语产生和听觉分割机制,基于离散音素序列构建计算模型,利用部分可观测马尔可夫决策过程,通过一系列假设探究言语障碍计算机制,为言语流畅性丧失研究及相关神经基质探索提供新途径。

中文摘要 AI 辅助

言语是复杂的运动和社交行为。本文介绍了基于离散音素序列的言语产生和听觉分割计算模型,目的是测试关于言语流畅性丧失机制的假设,该机制在口吃者或神经退行性疾病患者中会暂时或逐渐出现。模型采用部分可观测马尔可夫决策过程,核心特征包括词与不同数量音素序列的相互映射、言语生成和感知的共同状态子集,还有社会成分。提出了关于言语障碍计算机制的一系列假设,如降低下一个词的精度会导致口吃相关的停顿和词首重复,研究了这些假设预测的行为和信念更新模式,最后讨论了可能适用于功能成像研究的潜在神经基质。

英文摘要

Speech is a complex motor and social act. We must not only produce and understand sequences of variable-length words formed of ordered syllables but also speak and listen in turn. This theoretical paper introduces a computational model of speech production and auditory segmentation based upon sequences of discrete phonemes. The purpose of this is to develop a vehicle to test (in silico) hypotheses about the mechanisms that govern loss of speech fluency, something that may happen transiently in people who stutter or progressively in various neurodegenerative conditions. The modelling here appeals to Partially Observable Markov Decision Processes that provide a useful generic formulation for specifying the internal model our brains might use to describe dynamic environments in which they can exert partial control and make partial observations. The core features of the specific model used here are: (1) a reciprocal mapping from words to sequences of varying numbers of phonemes (2) a common subset of states for both speech generation and speech perception. A further novel feature is a social component, in which one must infer who is speaking (self or other). We propose a series of hypotheses about the computational mechanisms that might underwrite speech breakdown. For example, we demonstrate that reducing the precision in the next word introduces pauses and start-of-word repetitions of the sort associated with stuttering. We examine the patterns of behaviour and belief-updating predicted by each of these hypotheses, with implications for both speaking and listening. Finally, we discuss plausible neural substrates for the underlying message passing that might be amenable to interrogation with functional imaging.

补充信息

↑