arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ConvoDrift:用于建模文体语气演变的多轮对话数据集

ConvoDrift: A Multi-Turn Conversational Dataset for Modeling Stylistic Tone Evolution

Vihindi Kotalawala, Pamoda Dilranga, Gayani Thoradeniya, Prasan Yapa

arXiv 2610.02873首次发表:更新:

发表机构

Informatics Institute of Technology, Sri Lanka; National Health and Medical Research Council, Australia; Luxembourg Centre for Systems Biomedicine (LCSB), University of Luxembourg(斯里兰卡信息技术学院; 澳大利亚国家健康与医学研究理事会; 卢森堡大学系统生物医学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对对话中风格动态变化研究不足的问题,提出ConvoDrift数据集,含15,727个多轮结构及六组提示-响应对,支持风格漂移建模与个性化对齐研究,并通过人工与自动评估验证其有效性。

AI 中文摘要

对话中语言风格的演变是自然语言处理中一个尚未充分探索的问题。大多数风格控制数据集聚焦于句子层面,或假设整个对话过程中风格保持不变,从而忽略了用户偏好随交互变化而发生的动态转变。我们提出了ConvoDrift,一个旨在固定语义意图下对渐进式对话语气漂移进行建模的数据集。该数据集基于15,727个共享的多轮对话结构,用于适应性和人格条件对齐方法。每个对话包含六组提示-响应对,每组都标注了风格漂移和风格方向标签。这些对话对涵盖了多种沟通体裁。我们进一步通过配对语义等价但风格不同的响应,并使用五种不同的风格沟通人格来标注人格条件偏好,构建了一个互补的成对数据集,从而能够对语言语气中的个性化和多元对齐进行受控研究。除了数据集构建外,我们还进行了全面评估,包括人工验证、LLM作为评判者的评估以及自动词汇和语义评估。在由三位人工标注者标注的七项李克特标准中,平均Krippendorff's alpha为0.88,我们的词汇和语义分析表明,漂移事件引发词汇变化,同时保持语义相似性。

英文摘要

The evolution of linguistic style in conversations is an underexplored issue in NLP. Most style-control datasets focus on sentences or assume a static style throughout, missing the dynamic shifts that occur as user preferences change during interactions. We introduce ConvoDrift, a dataset designed to model progressive stylistic conversational tone drift under fixed semantic intent. It is built on 15,727 shared multi-turn conversational structures for adaptation and persona-conditioned alignment methods. It consists of six prompt-response pairs per conversation, each with the annotation of style drift and style direction labels. These pairs cover a range of communication genres. We further derive a complementary pairwise dataset by pairing semantically equivalent but stylistically distinct responses and annotating persona-conditioned preferences using five distinct style communication personas, enabling the controlled study of personalisation and pluralistic alignment in language tone. In addition to dataset construction, we conduct a comprehensive evaluation involving human validation, LLM-as-judge assessment, and automatic lexical and semantic evaluations. Across seven Likert criteria annotated by three human annotators, the average Krippendorff's alpha is 0.88, and our lexical and semantic analyses show that drift events induce lexical changes while preserving semantic similarity.

Comments13 pages, 14 figures, 7 tables, Accepted paper at the 13th Conference on Computational Linguistics and Speech Processing (ROCLING) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑