IndicTalk:用于印度语言的大规模基于角色的多语言对话语料库
IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
浏览论文内容
中文总结 AI 辅助
针对印度语言高质量多语言代码混合对话资源稀缺问题,提出通过全自动管道生成IndicTalk语料库,涵盖多种印度语言变体,经多种评估表现良好,将发布该语料库支持相关人工智能开发评估。
中文摘要 AI 辅助
大语言模型改变了对话式人工智能,但高质量的多语言代码混合对话资源仍然稀缺,尤其是对于印度语言。我们展示了IndicTalk,这是最大的多语言印度代码混合对话语料库之一,包含超过1328604个基于事件的多轮对话,涵盖9种印度语言的18种语言变体。该语料库通过全自动管道生成,结合了现实世界新闻基础、使用多语言大语言模型的角色条件对话生成和自动质量验证。广泛评估表明IndicTalk能产生流利、连贯且自然代码混合的对话。我们将发布IndicTalk以支持印度语言的多语言对话式人工智能的开发和评估。数据集可通过此https URL获取。
英文摘要
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .