用于对话策略优化的神经用户模拟器对抗学习
Adversarial learning of neural user simulators for dialogue policy optimisation
- Toshiba Europe Limited(东芝欧洲有限公司)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出用对抗学习训练神经用户模拟器,以生成更真实且多样的用户行为;餐厅搜索对话实验表明,其训练出的对话策略较最大似然模拟器成功率高8.3%。
AI中文摘要:
基于强化学习的对话策略通常是在与用户模拟器的交互中训练的。为了获得有效且鲁棒的策略,该模拟器应当生成既真实又多样的用户行为。当前的数据驱动模拟器经过训练,旨在准确建模对话语料库中的用户行为。我们提出一种使用对抗学习的替代方法,目标是模拟具有更多变化的真实用户行为。我们在餐厅搜索对话语料库上训练并评估了多个模拟器,随后使用它们训练对话系统策略。在策略交叉评估实验中,我们证明,与使用最大似然模拟器训练的策略相比,经过对抗训练的模拟器所产生的策略成功率高出8.3%。来自众包对话系统用户评估的主观结果证实了对抗训练用户模拟器的有效性。
英文摘要:
Reinforcement learning based dialogue policies are typically trained in interaction with a user simulator. To obtain an effective and robust policy, this simulator should generate user behaviour that is both realistic and varied. Current data-driven simulators are trained to accurately model the user behaviour in a dialogue corpus. We propose an alternative method using adversarial learning, with the aim to simulate realistic user behaviour with more variation. We train and evaluate several simulators on a corpus of restaurant search dialogues, and then use them to train dialogue system policies. In policy cross-evaluation experiments we demonstrate that an adversarially trained simulator produces policies with 8.3% higher success rate than those trained with a maximum likelihood simulator. Subjective results from a crowd-sourced dialogue system user evaluation confirm the effectiveness of adversarially training user simulators.