DuplexGen:人机交替对话的自适应合成
DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues
浏览论文内容
中文总结 AI 辅助
DuplexGen框架通过将LLM预测与少量槽级人类偏好标注校准,生成场景自适应交替发言对话,契合人类偏好的效果优于未校准方法,证明人类校准对交替发言合成场景特定性的关键作用。
中文摘要 AI 辅助
交替发言是全双工交互的核心组成部分,合适的交替发言行为会随场景变化,但当前模型无论上下文如何都采用单一规范。这一局限源于其训练数据:人-人语音语料库捕捉了自然时序现象,但几乎未提供角色基础或特定场景规范;而启发式或提示式合成方法注入交替发言行为时,未基于人类偏好。我们提出DuplexGen,这一框架通过将大语言模型(LLM)预测与少量槽级人类偏好标注校准,生成具有场景自适应交替发言的对话。在6项合作与竞争任务中,人类交替发言偏好存在系统性差异,DuplexGen与这些偏好的契合度远高于未校准提示或仅基于通用人-人数据的训练;在DuplexGen生成数据上训练的全双工模型表现出独特的、符合人类偏好的交替发言行为。这些结果表明,人类校准(而非仅语料库规模或提示设计)是交替发言合成实现场景特定性的关键。
英文摘要
Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models apply a single norm regardless of context. This limitation originates in their training data: human-human speech corpora capture natural timing phenomena but provide little role grounding or scenario-specific norms, while heuristic or prompted synthesis methods inject turn-taking behaviors without basing them on human preferences. We introduce DuplexGen, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations. In six cooperative and competitive tasks, human turn-taking preferences differ systematically, and DuplexGen aligns substantially more closely with those preferences than uncalibrated prompting or training solely on generic human-human data; a full-duplex model trained on DuplexGen-generated data exhibits distinctive, human-preferred turn-taking behaviors. These results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Seoul National University(首尔大学)
- Columbia University(哥伦比亚大学)
- University of California, Berkeley(加州大学伯克利分校)
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。