发表机构
AI Innovation Institute, Stony Brook University; University of Chicago; University of Illinois Urbana–Champaign; Social & Behavioral Science Institute, University of Illinois Urbana–Champaign(石溪大学人工智能创新研究院; 芝加哥大学; 伊利诺伊大学厄巴纳-香槟分校; 伊利诺伊大学厄巴纳-香槟分校社会与行为科学研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ISAAC是一个包含5.27亿条Reddit帖子的开放语料库,通过多步人工审核流程过滤,并标注语义标签,支持大规模分析社会群体话语的长期趋势与空间差异。
AI 中文摘要
我们介绍了伊利诺伊社会态度聚合语料库(ISAAC),这是一个开放、模块化且易于获取的语料库,包含超过5.27亿条英语Reddit帖子,这些帖子根据种族、性取向、年龄、能力、体重和肤色六个关键社会群体区分标准筛选而来,覆盖了从2007年到2023年的17年时间跨度。我们使用了一个多步骤、经过人工审核的过滤流程,以确保在整体以及每个社会群体区分标准下,策展数据集中不相关内容的比例保持在10%以下。随后,每条帖子通过算法标注了用户估计的家庭所在地区,以及一系列经过验证的现成和自定义语义标签,包括道德化、情感、情绪和语言概括性。我们通过汇聚证据验证了该语料库的有效性,将ISAAC与宏观社会趋势联系起来,例如在线搜索行为、重大社会事件(全国性和区域性)期间的时间性激增,以及公众态度的长期转变。通过提供统一、公开的基础设施,ISAAC消除了研究碎片化,并支持无缝复制,同时支持大规模多样化的实证工作流程。具体而言,ISAAC允许研究者进行跨类别比较,对社会群体话语的长期时间转变进行高精度追踪,并将空间差异映射到本地化的公众舆论和政策结果上。ISAAC完全公开、模块化的流程便于将语料库扩展到新的平台、语言和社会类别。为了适应各种研究需求,ISAAC既可以通过无需编码的点击式网站和标注者网络应用程序访问,也可以通过SQL游乐场、Python包和HuggingFace以编程方式访问。
英文摘要
We introduce the Illinois Social Attitudes Aggregate Corpus (ISAAC), an open, modular, and accessible corpus of 527 million+ English-language Reddit posts selected for relevance to six key social group distinctions based on race, sexuality, age, ability, body weight, and skin tone, covering the 17-year period from 2007 to 2023. A multi-step, human-audited filtering pipeline was used to keep irrelevant content in the curated dataset below 10%, both overall and for each social group distinction. Each post was then algorithmically annotated with the user's estimated home region, along with a suite of validated off-the-shelf and custom semantic labels including moralization, sentiment, emotion, and linguistic generalization. We confirm the validity of the resulting corpus through convergent evidence linking ISAAC to macro-level societal trends, such as online search behavior, temporal spikes during major societal events (both nationally and regionally), and long-term shifts in public attitudes. By offering a unified, public infrastructure, ISAAC eliminates research fragmentation and enables seamless replication while supporting diverse empirical workflows at scale. Specifically, ISAAC allows investigators to perform cross-category comparisons, conduct high-precision tracking of long-term temporal shifts in social group discourse, and map spatial variation onto localized public opinion and policy outcomes. ISAAC's fully public, modular pipeline facilitates easy extension of the corpus to new platforms, languages, and social categories. To accommodate various research needs, ISAAC is accessible both without coding through a point-and-click website and labeler web-apps, and programmatically via an SQL playground, a Python package, and HuggingFace.
CommentsSubmitted to Behavior Research Methods