发表机构
Data Science and Innovation Ata Technology Platforms; Digital Operations and Data(阿塔技术平台数据科学与创新部; 数字运营与数据部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建多模态土耳其语对话数据集,采用遗传算法优化规则,实现话轮转换预测,填补土耳其语相关自然对话语料库的空白。
AI 中文摘要
话轮转换是人类对话的基本组织特征,在自然同步对话系统中仍难以建模。现有研究已探索多模态方法和大语言模型用于话轮结束预测,但缺乏专门针对土耳其语话轮转换动态的自然对话语料库。本研究引入一个多模态土耳其语对话数据集,包含无剧本的双人交互,涵盖同步正面视频、可将重叠语音归因于个体说话者的分说话者音频通道,以及时间对齐的转录文本。将话轮转换预测视为二分类问题,采用遗传算法(GA)优化从视觉、声学和语言特征中推导的可解释决策规则,所提框架采用混合AND-OR规则表示,以代表话轮转换前的替代线索组合。
英文摘要
Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems. While existing research has explored multimodal approaches and large language models for turn-ending prediction, there is a lack of naturalistic conversational corpora specifically addressing turn-taking dynamics in Turkish. This study introduces a multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions. Turn-taking prediction is formulated as a binary classification problem, and a Genetic Algorithm (GA) is employed to optimize interpretable decision rules derived from visual, acoustic, and linguistic features. A hybrid AND-OR rule representation is adopted in the proposed framework to represent the alternative cue combinations that precede a turn transition.
CommentsAccepted to INTCEC 2026. This is the author's pre-print version. The final authenticated version will be available through the conference proceedings