arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型了解其俚语吗?用户生成内容中的酷儿俚语理解

Do Language Models Know Their Slang? Queer Slang Understanding in User-Generated Content

Arianna Denitto, Beatrice Savoldi

arXiv 2608.04847首次发表:更新:

发表机构

University of Torino; Fondazione Bruno Kessler(都灵大学; 布鲁诺·凯塞勒基金会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对酷儿俚语在自然语言处理研究中未受充分关注的问题,构建含118个酷儿术语的Slang-Q数据集,对语言模型理解酷儿俚语的能力开展首次探索性评估,为研究模型处理敏感社区语言的能力提供基础。

AI 中文摘要

尽管酷儿俚语具有文化相关性且传播广泛,但它在自然语言处理研究中仍未得到充分体现。为解决这一差距,我们推出Slang-Q,这是一个人工整理的数据集,包含自然用户生成的英文句子,配对有酷儿俚语术语和参考定义,该数据集基于新构建的包含118个酷儿术语的分类体系构建。我们利用该资源对语言模型在不同提示条件下理解和定义酷儿俚语的能力进行首次探索性评估。Slang-Q旨在作为研究当前模型如何处理敏感的、特定社区的语言,以及它们是否能提供关于此类身份和语言表达形式的准确可靠信息的基础。

英文摘要

Despite its cultural relevance and diffusion, queer slang remains underrepresented in Natural Language Processing research. Towards addressing this gap, we introduce Slang-Q, a manually curated dataset of naturally user-generated English sentences paired with queer slang terms and reference definitions, built upon a newly constructed taxonomy of 118 queer terms. We use this resource to conduct a first exploratory evaluation of language models on their ability to understand and define queer slang under varying prompting conditions. Slang-Q is intended as a basis for studying how current models handle sensitive, community-specific language and whether they can provide accurate and reliable information about such forms of identity and linguistic expression.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑