CultureTalk-ID:印尼当地语言文化常识的多任务对话基准测试
CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages
查看机构详情
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对印尼语文化常识基准测试忽略对话语境的问题,引入CultureTalk-ID基准测试,涵盖多种语言和主题,通过多阶段人工流程确保真实性,并引入三个互补任务,探究语言模型在文化常识方面的理解、转换和生成能力。
中文摘要 AI 辅助
文化通过对话得以体现,但现有的印尼文化常识基准测试在短且孤立的提示上评估语言模型,忽略了文化细微差别得以呈现的对话语境。我们引入了CultureTalk-ID,这是首个针对印尼语及其当地语言文化常识的基于对话的基准测试,包含4496个跨11种语言和13个文化突出主题的有文化根基的对话,通过多阶段人工流程与母语人士策划以确保真实性。CultureTalk-ID引入了三个互补任务,共同探究语言模型是否能理解、转换和生成有文化根基的语言。
英文摘要
Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping away the dialogic context in which cultural nuances actually surface. We introduce CultureTalk-ID, the first dialogue-based benchmark for cultural commonsense in Indonesian and its local languages, comprising 4,496 culturally grounded dialogues across 11 languages and 13 culturally salient topics, curated through a multi-stage human pipeline with native speakers to ensure authenticity. CultureTalk-ID introduces three complementary tasks, namely dialogue-based multiple-choice cultural commonsense reasoning, culturally faithful machine translation, and language steering, which jointly probe whether LLMs can understand, transfer, and generate culturally grounded language.