arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12894cs.CL

BavGround:巴伐利亚区域文化接地与方言能力基准

BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian

Jophin John, Michael Hoffmann, Jan Fillies, Michael A. Hedderich, Barbara Plank

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出BavGround基准,评估LLM的巴伐利亚区域文化接地与方言能力,发现多语言模型在巴伐利亚语及源接地问题上表现欠佳,且评估协议会显著影响结论。

中文摘要 AI 辅助

对大型语言模型(LLM)的文化评估通常聚焦于高资源标准语言,导致区域文化和方言社区代表性不足。我们推出BavGround,这是一个用于评估英语、德语和巴伐利亚语的巴伐利亚区域文化接地与方言能力的基准。BavGround包含每种语言8个文化领域的206道多项选择题源问题,共产生618个多并行实例,题目涵盖广泛可获取的文化知识以及来自新闻、历史资料和专业文献的源接地区域知识。我们评估了15个7B-10B规模的开放权重指令调优模型和1个闭源模型参考。强多语言模型总体表现最佳,但在巴伐利亚语题目和源接地问题上表现下降,表明其在方言和本地化文化知识方面仍存在困难。我们进一步表明,结论在很大程度上取决于评估协议:原始答案字母评分、打乱字母评分、选项文本似然、生成答案解析和语义匹配会产生不同的绝对分数和排名,尤其对于区域适配模型。最后,对GENBA-10B检查点的探索性分析表明,持续预训练在各领域对答案内容似然的提升不均衡,而方言能力仍相对较弱。BavGround支持对LLM文化表征的本地化、协议感知评估。

英文摘要

Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian regional cultural grounding and dialect competence across English, German and Bavarian. BavGround contains 206 multiple-choice source questions across eight cultural domains per language, yielding 618 multi-parallel instances, with items covering both broadly accessible cultural knowledge and source-grounded regional knowledge from journalism, historical sources, and specialist literature. We evaluate fifteen 7B-10B open-weight instruction-tuned models and one closed-model reference. Strong multilingual models perform best overall, but performance drops on Bavarian items and source-grounded questions, indicating persistent difficulty with dialectal and localized cultural knowledge. We further show that conclusions depend strongly on evaluation protocol: raw answer-letter scoring, shuffled-letter scoring, option-text likelihood, generated-answer parsing, and semantic matching can produce different absolute scores and rankings, especially for regionally adapted models. Finally, an exploratory analysis of GENBA-10B checkpoints suggests that continued pretraining improves answer-content likelihood unevenly across domains, while dialect competence remains comparatively weak. BavGround supports localized, protocol-aware evaluation of cultural representation in LLMs.

发表机构

  • Stanford University(斯坦福大学)
  • Freie Universität Berlin(柏林自由大学)
  • Center for Information and Language Processing, LMU Munich(慕尼黑大学信息与语言处理中心)
  • Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
  • Leibniz Supercomputing Centre (LRZ)(莱布尼茨超级计算中心)

机构由 AI 辅助整理,请以论文原文为准。

↑