arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03321cs.CL

将对话轮次与语义解耦:一种基于有限状态机的全双工对话的解耦数据方法

Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex Dialogue

  • Kyoto University(京都大学)

机构由 AI 辅助整理,请以论文原文为准。

Yihang Li, Chenhui Chu

AI总结:

该研究针对神经有限状态机框架依赖合成数据导致对话轮次自然度不足的问题,提出解耦数据方法,通过规则转换HH对话为FSM序列并结合SAC损失,提升对话轮次能力同时恢复LLM语义能力。

AI中文摘要:

神经有限状态机(NFSM)框架通过在标准的下一个token预测目标下将对话轮次控制和响应生成序列化到单个因果序列上,以较低的微调成本保留了语义能力,为全双工对话提供了一种实用路径。然而,其对合成文本数据的依赖从根本上限制了对话轮次的自然度,因为大型语言模型(LLM)无法忠实地模拟真实人类对话的细粒度声学时间动态。在这项工作中,我们提出了一种解耦数据方法,该方法从真实的人机(HH)口语对话中学习对话轮次,同时通过可配置的人机(HA)文本对话塑造语义行为。为了实施该方法,我们引入了一种基于规则的事件引导数据转换方法,该方法通过对对话轮次事件进行分类并应用确定性映射规则,将HH口语对话序列化为有限状态机(FSM)序列,从而在无需LLM生成的注释的情况下实现可扩展的监督。我们进一步提出了一种源感知校准(SAC)损失,该损失共同校准状态转换token的长尾分布,并将每个数据源导向其最适合监督的能力。实验表明,我们的方法显著提高了对话轮次的熟练度,同时恢复了基础LLM的语义能力。我们的代码和模型可在此https URL获取。

英文摘要:

The Neural Finite State Machine (NFSM) framework offers a pragmatic path to full-duplex dialogue by serializing turn-taking control and response generation onto a single causal tape under the standard next-token prediction objective, thereby preserving semantic prowess at a low fine-tuning cost. However, its reliance on synthetic text data fundamentally limits turn-taking naturalness, as Large Language Models (LLMs) cannot faithfully simulate the fine-grained acoustic temporal dynamics of real human dialogues. In this work, we propose a decoupled data approach that learns turn-taking from real Human-Human (HH) spoken dialogues while shaping semantic behavior through configurable Human-Agent (HA) text dialogues. To operationalize this approach, we introduce a rule-based event-guided data transformation method that serializes HH spoken dialogues into FSM tapes by classifying turn-taking events and applying deterministic mapping rules, enabling scalable supervision without LLM-generated annotations. We further propose a Source-Aware Calibrated (SAC) Loss that jointly calibrates the long-tailed distribution of state transition tokens and channels each data source toward the capability it best supervises. Experiments show that our approach substantially improves turn-taking proficiency while recovering the foundation LLM's semantic capability. Our code and model are available at https://github.com/Liyht/def-fsm.

补充信息

↑