arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

嵌入多样性:一种基于嵌入的数据多样性测量工具

emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity

Cantao Su, Menan Velayuthan, Esther Ploeger, Dong Nguyen, Anna Wegmann

arXiv 2607.19848首次发表:更新:

发表机构

Department of Information & Computing Sciences, Utrecht University(乌得勒支大学信息与计算科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在解决缺乏基于嵌入量化数据多样性标准化工具的问题,核心方法是提供emb-diversity工具,主要贡献是展示该工具在测量数据集多种多样性方面的潜力。

AI 中文摘要

越来越多证据表明数据多样性对开发公平且强大的自然语言处理模型至关重要。但当前测量多样性的方法不一致且零散,缺乏基于嵌入量化多样性的标准化工具。基于嵌入的多样性度量很灵活,可用于任何嵌入模型和可嵌入数据。我们提供了基于嵌入的综合多样性测量工具emb-diversity,展示了其在多种用例中的潜力,如测量数据集的文体、语义、语言和说话者多样性。

英文摘要

There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmented: While there exist a number of tools for measuring the lexical diversity of texts, researchers lack standardized tools for quantifying diversity based on embeddings. Embedding-based diversity measures are highly flexible: They work with any embedding model and any data that can be embedded, and are thus applicable to many notions of diversity. With emb-diversity, we provide a comprehensive embedding-based diversity measurement tool, spanning a broad range of measures. We demonstrate its potential for several use cases: measuring the stylistic, semantic, language and speaker diversity of datasets. https://github.com/nlpsoc/emb-diversity/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑