发表机构
Graduate School of Informatics, Kyoto University(京都大学信息学研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多方对话中话语 addressed 对象的识别问题,通过构建二元地址标签和连续地址层级,研究其与话轮转换及听众行为的关系,发现连续地址层级模型预测性更好,为受话人检测研究指明未来方向。
AI 中文摘要
在对话系统与多用户的多方对话中,识别话语的 addressed 对象是关键挑战。以往工作将其视为多类分类任务,假定地址是离散的,主要用于预测话轮转换。本文通过将地址视为连续现象重新审视该假设。利用多方人工对话语料库,构建了基于多数投票的二元地址标签和通过潜在变量模型从注释者判断推断出的连续地址层级。接着研究这些表示与话轮转换及听众行为(如注视和反馈渠道)的关系。结果表明,除话轮转换外,注视和反馈渠道也与地址相关。使用连续地址层级的模型比使用离散标签的模型具有更好的预测拟合度,这表明地址可能呈现分级结构。最后基于研究结果讨论了受话人检测研究的未来方向。
英文摘要
In multi-party dialogues between a dialogue system and multiple users, identifying to whom an utterance is addressed is a key challenge. Prior work has typically treated addressee detection as a multi-class classification task, selecting a single label representing an individual participant or the group. This formulation assumes that address is inherently discrete and has primarily been used for predicting turn-taking. In this paper, we revisit this assumption by analyzing address as a continuous phenomenon. Using a multi-party human dialogue corpus annotated by multiple annotators, we construct both binary address labels derived from majority-vote addressee labels and continuous address levels inferred from annotator judgments using a latent-variable model. We then examine how these representations relate to turn-taking as well as listener behaviors, including gaze and backchannels. Our results show that, in addition to turn-taking, both gaze and backchannels are associated with address. Furthermore, models using continuous address levels achieve better predictive fit than those using discrete labels, suggesting that address may exhibit graded structure. Finally, we discuss the future directions of addressee detection research based on the findings of this study.