arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CultureMINE:提升NLP系统文化能力的数据集与方法

CultureMINE: Datasets and Methods for Improving the Cultural Capabilities of NLP Systems

Tania Chakraborty, Eylon Caplan, Zhaoqing Wu, Kevin Cushing, Han Qin, Shreya Havaldar, Dan Goldwasser

arXiv 2609.22494首次发表:更新:

发表机构

Purdue University; University of Pennsylvania(普渡大学; 宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过分析375余篇论文,系统梳理文化NLP中的能力目标、数据创建与方法趋势,并发布交互式论文列表以促进社区研究。

AI 中文摘要

近年来,文化NLP(Cultural NLP)领域引起了广泛关注,人们投入了大量努力来构建全球包容性的NLP系统。该领域文献的快速增长使得追踪方法和数据资源的趋势变得困难。为了解决这一问题,我们分析了超过375篇论文,以回答三个互补的问题:(1)NLP系统中针对哪些文化能力(Cultural Capabilities, CCs)?(2)文化数据资源是如何创建的?(3)使用哪些方法来提升这些系统的文化能力?我们讨论了在这三个问题中观察到的趋势,并指出了相关的研究空白。为了促进该领域的进一步研究,我们以交互式网页界面的形式发布了我们分析的全部论文列表,该界面包含一项功能,允许研究人员添加他们的工作;我们希望这能促进未来的研究,并成为文化NLP社区的一项宝贵资源。

英文摘要

In recent years, there has been a surge of interest in Cultural NLP, with substantial efforts to create globally inclusive NLP systems. The rapid growth of literature in this field makes it difficult to track trends in methods and data resources. To address this, we analyze over 375 papers to answer three complementary questions: (1) What Cultural Capabilities (CCs) are being targeted in NLP systems? (2) How are cultural data resources being created? and (3) What methods are being used to improve the CCs of those systems? We discuss trends observed across the three questions, and identify relevant research gaps. To facilitate further research in this field, we release our full list of analyzed papers in the form of an interactive web interface, which includes a feature to allow researchers to add their work; we hope this facilitates future research and proves to be a valuable resource for the Cultural NLP community.

CommentsAccepted to NLP+CSS 2026

Journal refProceedings of the Seventh Workshop on Natural Language Processing and Computational Social Science (2026), pp. 198-248

DOI:10.18653/v1/2026.nlpcss-1.14

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑