arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31264cs.CL

识别X平台上的科学家

Identifying Scientists on X

  • Heinrich-Heine-University(海因里希·海涅大学)
  • GESIS - Leibniz Institute for the Social Sciences(GESIS 莱布尼茨社会科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

Philipp Meier, Katarina Boland, Laura Kallmeyer, Stefan Dietze

AI总结:

本研究提出基于用户简介和推文自动识别X平台科学家用户的方法,随机森林F1达0.88,集成DeBERTa模型达0.96,并提供两个标注数据集。

AI中文摘要:

随着科学相关话语在网络上的重要性日益增长以及传统知识秩序的侵蚀,自动识别不同用户群体(如科学家)变得至关重要。本研究提出了一种基于用户个人简介和推文来识别X/Twitter上科学家与非科学家用户的方法。我们展示了在两个不同数据集上,使用带有语言特征的随机森林模型能够将账户分类为科学家和非科学家,F1分数最高达到0.88;而使用对比微调的DeBERTa模型在集成设置中,F1分数最高达到0.96。此外,我们提供了两个数据集,其中包含被标记为科学家或非科学家的X用户及其各自的推文和用户个人简介。

英文摘要:

With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientists and non- scientists on X/Twitter based on their user biographies and tweets. We show that we are able to classify accounts as scientists and non- scientists on two different datasets, reaching an F1 score of up to 0.88 using Random Forests with linguistic features and up to 0.96 using a contrastively fine-tuned DeBERTa model in an ensemble setup. Furthermore, we provide two datasets with X users labeled as scientists or non scientists and their respective tweets and user biographies.

补充信息

↑