arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2401.04867cs.CLcs.AIcs.HC

用于客观评估口语对话系统的用户行为分析

An Analysis of User Behaviors for Objectively Evaluating Spoken Dialogue Systems

  • Kyoto University(京都大学)
  • KTH Royal Institute of Technology(瑞典皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Koji Inoue, Divesh Lala, Keiko Ochi, Tatsuya Kawahara, Gabriel Skantze

更新

AI总结:

本文提出基于用户行为客观评估口语对话系统的框架,分析专注倾听、求职面试和初次见面交谈中的行为指标与主观评分关系,揭示不同任务应采用不同行为指标。

AI中文摘要:

建立口语对话系统的评估方案十分重要,但也可能具有挑战性。虽然主观评估常用于用户实验,但客观评估对于研究比较和可复现性是必要的。为解决这一问题,我们提出了一个基于用户行为间接但客观评估系统的框架。为此,本文研究了社交对话任务中用户行为与主观评估分数之间的关系,这些任务包括专注倾听、求职面试和初次见面交谈。结果表明,在以用户话语为主的对话任务中,如专注倾听和求职面试,话语数量和词语数量等指标在评估中发挥重要作用。观察不流利现象也可以指示正式任务(如求职面试)的有效性。另一方面,在互动性较高的对话任务中,如初次见面交谈,与话轮转换相关的行为(如平均切换停顿时长)变得更加重要。这些发现表明,选择适当的用户行为可以为每项社交对话任务中的客观评估提供有价值的见解。

英文摘要:

Establishing evaluation schemes for spoken dialogue systems is important, but it can also be challenging. While subjective evaluations are commonly used in user experiments, objective evaluations are necessary for research comparison and reproducibility. To address this issue, we propose a framework for indirectly but objectively evaluating systems based on users' behaviors. In this paper, to this end, we investigate the relationship between user behaviors and subjective evaluation scores in social dialogue tasks: attentive listening, job interview, and first-meeting conversation. The results reveal that in dialogue tasks where user utterances are primary, such as attentive listening and job interview, indicators like the number of utterances and words play a significant role in evaluation. Observing disfluency also can indicate the effectiveness of formal tasks, such as job interview. On the other hand, in dialogue tasks with high interactivity, such as first-meeting conversation, behaviors related to turn-taking, like average switch pause length, become more important. These findings suggest that selecting appropriate user behaviors can provide valuable insights for objective evaluation in each social dialogue task.

补充信息

↑