BavGround:巴伐利亚区域文化接地与方言能力基准
BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian
浏览论文内容
中文总结 AI 辅助
该研究推出BavGround基准,评估LLM的巴伐利亚区域文化接地与方言能力,发现多语言模型在巴伐利亚语及源接地问题上表现欠佳,且评估协议会显著影响结论。
中文摘要 AI 辅助
对大型语言模型(LLM)的文化评估通常聚焦于高资源标准语言,导致区域文化和方言社区代表性不足。我们推出BavGround,这是一个用于评估英语、德语和巴伐利亚语的巴伐利亚区域文化接地与方言能力的基准。BavGround包含每种语言8个文化领域的206道多项选择题源问题,共产生618个多并行实例,题目涵盖广泛可获取的文化知识以及来自新闻、历史资料和专业文献的源接地区域知识。我们评估了15个7B-10B规模的开放权重指令调优模型和1个闭源模型参考。强多语言模型总体表现最佳,但在巴伐利亚语题目和源接地问题上表现下降,表明其在方言和本地化文化知识方面仍存在困难。我们进一步表明,结论在很大程度上取决于评估协议:原始答案字母评分、打乱字母评分、选项文本似然、生成答案解析和语义匹配会产生不同的绝对分数和排名,尤其对于区域适配模型。最后,对GENBA-10B检查点的探索性分析表明,持续预训练在各领域对答案内容似然的提升不均衡,而方言能力仍相对较弱。BavGround支持对LLM文化表征的本地化、协议感知评估。
英文摘要
Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introduce BavGround, a benchmark for evaluating Bavarian regional cultural grounding and dialect competence across English, German and Bavarian. BavGround contains 206 multiple-choice source questions across eight cultural domains per language, yielding 618 multi-parallel instances, with items covering both broadly accessible cultural knowledge and source-grounded regional knowledge from journalism, historical sources, and specialist literature. We evaluate fifteen 7B-10B open-weight instruction-tuned models and one closed-model reference. Strong multilingual models perform best overall, but performance drops on Bavarian items and source-grounded questions, indicating persistent difficulty with dialectal and localized cultural knowledge. We further show that conclusions depend strongly on evaluation protocol: raw answer-letter scoring, shuffled-letter scoring, option-text likelihood, generated-answer parsing, and semantic matching can produce different absolute scores and rankings, especially for regionally adapted models. Finally, an exploratory analysis of GENBA-10B checkpoints suggests that continued pretraining improves answer-content likelihood unevenly across domains, while dialect competence remains comparatively weak. BavGround supports localized, protocol-aware evaluation of cultural representation in LLMs.
发表机构
- Stanford University(斯坦福大学)
- Freie Universität Berlin(柏林自由大学)
- Center for Information and Language Processing, LMU Munich(慕尼黑大学信息与语言处理中心)
- Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
- Leibniz Supercomputing Centre (LRZ)(莱布尼茨超级计算中心)
机构由 AI 辅助整理,请以论文原文为准。