AI 中文总结
研究 GitHub 毒性标注问题,提出人工参与注释方法,通过小型本地语言模型预测及随机森林验证器筛选,降低标注成本,应用此流程处理大量对话,评估先前研究发现并给出新见解。
AI 中文摘要
开源讨论中的有害互动会疏远贡献者并威胁项目可持续性,但此前对 GitHub 毒性的实证研究规模有限,其普遍性存疑。扩大规模有困难,因为 GitHub 上的毒性往往是隐含且依赖上下文的。我们提出一种人工参与(HITL)注释方法,使大规模、领域校准的毒性标签实用化。单次调用小型本地语言模型可生成毒性预测和一组可解释的事件类别分数。轻量级随机森林验证器利用这些分数标记最可能被误标记的小部分对话,仅在需要时引导人工审查。验证器优于基于置信度和多语言模型的基线,同时注释成本低。我们将此流程应用于超过 124,000 个 GitHub 问题和拉取请求对话。利用所得数据集,我们评估了先前小规模研究的关键发现,证实了一些并修正了其他一些,还对不同开源项目中毒性的普遍性、特征和动态提出了新见解。
英文摘要
Toxic interactions in open source discussions can alienate contributors and threaten project sustainability, yet prior empirical studies of GitHub toxicity have been limited in scale, raising questions about their generalizability. Scaling up is difficult because toxicity on GitHub is often implicit and context-dependent, making both fully manual annotation and LLM-based labeling unreliable. We present a human-in-the-loop (HITL) annotation methodology that makes large-scale, domain-calibrated toxicity labeling practical. A single call to a small, local LLM produces both a toxicity prediction and a set of interpretable event category scores. A lightweight Random Forest validator then uses those scores to flag the small subset of conversations most likely to be mislabeled, directing human review only where it is needed. The validator outperforms confidence-based and multi-LLM baselines while adding low annotation cost. We apply this pipeline to over 124,000 GitHub issue and pull request conversations. Using the resulting dataset, we evaluate key findings from prior small-scale research, confirming some and qualifying others, and present new insights into the prevalence, characteristics, and dynamics of toxicity across diverse open source projects.
CommentsAccepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)