arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你构建框架:概念表征如何塑造大语言模型对反犹主义的检测与推理

You Frame It: How Conceptual Representations Shape LLM Detection and Reasoning about Antisemitism

Katharina Soemer, Helena Mihaljević

arXiv 2607.04945首次发表:更新:

发表机构

Goethe University, Frankfurt; HTW Berlin(法兰克福歌德大学; 柏林工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究不同形式概念基础对四个先进大语言模型检测反犹主义及解释行为的影响,用两个专家注释数据集比较不同表征,发现细粒度分类表征能提高召回率但降低精度,还揭示了模型解释存在的局限。

AI 中文摘要

大语言模型在推理时能整合外部概念资源,为检测反犹主义等复杂现象创造新机会。我们研究了不同形式的概念基础如何影响四个最先进大语言模型的反犹主义检测和解释行为。使用两个专家注释数据集,我们比较了反犹主义的定义、细粒度分类、示例增强和大语境表征。我们发现细粒度分类表征显著提高了召回率,同时降低了精度。令人惊讶的是,提供大量更大的概念资源并没有带来额外的定量好处。大屠杀后的反犹主义在所有模型和配置中都构成了最持久的挑战。对解释的分析进一步揭示了系统的局限性,包括概念参考的过度产生、对词汇线索的依赖、过度自信以及处理微妙或正当形式的反犹主义的困难。我们的发现突出了基于概念的大语言模型在反犹主义检测和推理方面的潜力和剩余局限性。

英文摘要

LLMs enable the integration of external conceptual resources at inference time, creating new opportunities for detecting ideologically and historically complex phenomena such as antisemitism. We investigate how different forms of conceptual grounding affect antisemitism detection and explanation behavior across four state-of-the-art LLMs. Using two expert-annotated datasets, we compare definitional, fine-grained taxonomic, example-augmented, and large-context representations of antisemitism. We find that fine-grained taxonomic representations substantially improve recall, while simultaneously reducing precision. Surprisingly, supplying substantially larger conceptual resources yields no additional quantitative benefit. Post-Holocaust antisemitism poses the most persistent challenge across models and configurations. Analysis of explanations further reveals systematic limitations including overproduction of conceptual references, reliance on lexical cues, overconfidence, and difficulties with subtle or justificatory forms of antisemitism. Our findings highlight both the potential and the remaining limitations of conceptually grounded LLMs for antisemitism detection and reasoning.

CommentsAccepted to Findings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑