arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TalkFa:面向波斯语对话生成与理解的统一基准

TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding

Neda Jamshidi, Kamyar Zeinalipour, Fahimeh Akbari, Monica Bianchini, Marco Maggini, Marco Gori

arXiv 2609.01810首次发表:更新:

发表机构

University of Siena(锡耶纳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对波斯语缺乏对话生成与理解综合基准的问题,提出含三个互补数据集的TalkFa基准,经实验验证其可靠性,还将发布相关资源。

AI 中文摘要

波斯语使用者超过1.2亿人,但目前缺乏针对对话生成与理解的综合基准。本文介绍TALKFA,这一统一基准包含三个互补数据集:(1)WIKI-FADIAL,共4200条基于维基百科的对话,用于知识驱动生成;(2)DAILYDIALOG-FA,共6600条标注了对话行为与情感的对话;(3)PLAYDIAL-FA,共2100条带有情感标签的戏剧对话。尽管大型语言模型(LLMs)可辅助数据构建,但每条对话都经过波斯语母语者的多轮审核与修订,仅最终经人类批准的对话会被发布。对6个LLAMA和MISTRAL模型的实验表明,LoRA能显著提升对话生成效果,且仅需25%-50%的训练数据即可恢复90%以上的最终性能提升。在分类任务中,FABERT取得最佳对话行为性能,LoRA-MISTRAL-7B在情感识别上表现最优,MISTRAL-24B获得最高情感得分。人工评估与独立外部验证证实了该基准的可靠性,与作为LLM评判的GPT-4.1的对比显示,自动指标会大幅高估对话质量。对前沿LLMs的零样本评估进一步表明,TalkFa仍是一项具有挑战性的基准。我们将发布所有数据集、标注指南、代码及检查点。

英文摘要

Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding. We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-grounded generation; (2) DAILYDIALOG-FA, 6.6K dialogues annotated for dialogue acts and emotions; and (3) PLAYDIAL-FA, 2.1K theatrical dialogues with sentiment labels. While LLMs assist data construction, every dialogue undergoes multi-stage review and revision by native Farsi speakers, and only the final human-approved dialogues are released. Experiments with six LLAMA and MISTRAL models show that LoRA substantially improves dialogue generation while requiring only 25-50% of the training data to recover over 90% of the final performance gains. Across classification tasks, FABERT achieves the best dialogue-act performance, LORA-MISTRAL-7B performs best on emotion recognition, and MISTRAL-24B achieves the highest sentiment score. Human evaluation and independent external validation demonstrate the reliability of the benchmark, while comparisons with GPT-4.1 as an LLM judge reveal that automatic metrics substantially overestimate dialogue quality. Zero-shot evaluation with frontier LLMs further shows that TalkFa remains a challenging benchmark. We will release all datasets, annotation guidelines, code, and checkpoints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑