发表机构
Shanghai Jiao Tong University; Shanghai Innovation Institute; Shanghai AI Laboratory; MOSI Intelligence; SenseTime Group Inc.; Shenzhen University of Advanced Technology(上海交通大学; 上海创新研究院; 上海人工智能实验室; MOSI Intelligence; 商汤科技集团; 深圳先进技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
XTurnix通过双状态二元决策实现统一话轮控制,预训练于550万因果动作示例,在多个基准上优于基线,准确率达89.06%。
AI 中文摘要
实时对话系统中的一般性话轮转换行为需要在听讲时决定是继续聆听还是开始回应,以及在说话时决定是继续还是停止。现有的话轮检测器使用异构的、任务特定的标签空间,并且通常在有限的标注数据上训练或在孤立的语句上评估,这使得它们难以作为具有全面上下文的统一因果控制器使用。我们提出了XTurnix,一个紧凑的基于文本的模型,它将话轮控制形式化为基于AI当前听讲或说话状态的两个二元决策,并从完整的对话历史中预测单个控制标记。XTurnix在从带时间戳的双说话者转录中自动派生的550万个因果动作示例上进行预训练,然后在具有四个状态-动作标签更平坦分布的合成多轮示例上进行微调。我们在四个公共基准和一个平衡的自策基准上评估XTurnix。在公共基准上,XTurnix在所有SemanticVAD和LiveKit分割上取得了最佳结果,在Smart-Turn Bench上与原生Smart-Turn模型持平,并在Easy-Turn上取得了最高的不完整话轮准确率。在自策基准上,它达到了89.06%的准确率,比最强的第三方基线(68.75%)高出20多个百分点,同时在所有四个类别中保持了84.21%至90.63%的F1分数。这些结果表明,在一个紧凑模型中实现了统一的听讲和说话状态话轮控制。代码可在本https URL获取,交互式演示可在本https URL获取。
英文摘要
General turn-taking behavior in real-time dialogue systems requires deciding whether to keep listening or start responding while listening, and whether to continue or stop while speaking. Existing turn detectors use heterogeneous, task-specific label spaces and are often trained on limited annotations or evaluated on isolated utterances, making them difficult to use as a unified causal controller with comprehensive context. We propose XTurnix, a compact text-based model that formulates turn control as two binary decisions conditioned on the AI's current listening or speaking state and predicts a single control token from the complete dialogue history. XTurnix is pretrained on 5.5 million causal action examples automatically derived from timestamped two-speaker transcripts, then fine-tuned on synthetic multi-turn examples with a flatter distribution across the four state-action labels. We evaluate XTurnix on four public benchmarks and a balanced self-curated benchmark. Across the public benchmarks, XTurnix achieves the best results on all SemanticVAD and LiveKit splits, ties the native Smart-Turn model on Smart-Turn Bench, and achieves the highest incomplete-turn accuracy on Easy-Turn. On the self-curated benchmark, it reaches 89.06% accuracy, more than 20 percentage points above the strongest third-party baseline at 68.75%, while maintaining F1 scores between 84.21% and 90.63% across all four categories. These results demonstrate unified listening- and speaking-state turn control in a single compact model. Code is available at https://github.com/xcc-zach/xturnix, with an interactive demo at https://huggingface.co/spaces/xcczach/xturnix-demo.