arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30897cs.CL

从标注到推理:语言模型中的文化

From annotation to reasoning: Culture in language models

  • University of Copenhagen(哥本哈根大学)

机构由 AI 辅助整理,请以论文原文为准。

Daniel Hershcovich, Alexander Conroy, Jens Bjerring-Hansen

AI总结:

针对文化基准仅测事实或认同的局限,提出以文学解释为场景,通过基于证据的基准、保留分歧的评估和模型开发实验,推动AI文化鲁棒性。

AI中文摘要:

当不止一种解释可能正确时,我们应如何评估语言模型?文化基准通常测试事实知识、对调查回答的认同或对预定义含义的识别。这些任务未涉及模型能否解释文化引用在特定文本中如何运作、用证据支持某种解读或在受到批评后修正解读。这是一个解释深度的问题,与文化覆盖的广度互补。我们认为文学解释为研究这些能力提供了有用的场景。我们聚焦于文化引用与复用:文本如何在历史和语言语境中援引、重复和改造早期表达。我们的核心主张是,文学学者可能对某种解释存在分歧,同时认可其支持的品质。我们提出将基于证据的基准、保留学者分歧的评估,以及关于文学数据、语境资源和学者反馈的模型开发实验联系起来。丹麦文学提供了一个具体的起点,对其他语言和领域具有启示意义。目标是开发超越传统基准指标的替代评估策略,并引导模型开发走向AI系统中的文化鲁棒性。

英文摘要:

How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain how a cultural reference works in a particular text, support a reading with evidence, or revise it after criticism. This is a question of interpretive depth, complementary to the breadth of cultural coverage. We argue that literary interpretation offers a useful setting for studying these capabilities. We focus on cultural referencing and reuse: how texts invoke, repeat, and transform earlier expressions across historical and linguistic contexts. Our central claim is that literary scholars can disagree about an interpretation while recognizing the quality of its support. We propose linking evidence-centered benchmarks, evaluation that preserves scholarly disagreement, and model-development experiments on literary data, contextual resources, and scholarly feedback. Danish literature provides a concrete starting point, with implications for other languages and domains. The aim is to develop alternative evaluation strategies that go beyond conventional benchmark metrics and guide model development toward cultural robustness in AI systems.

补充信息

↑