发表机构
George Mason University(乔治梅森大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对自闭症患者写作常被误判为AI生成的传闻,通过约60000篇Reddit帖子语料库,对比OpenAI GPT - 2检测模型输出概率,发现自闭症子语料库被误判更多,二者特征联系不直接,呼吁对相关模型及其学术应用进行伦理审查。
AI 中文摘要
近期研究表明,人工智能检测模型无法准确识别人工智能生成的文本,且可能对某些少数群体存在偏见。本研究通过实证检验了关于自闭症作家的作品更常被标记为人工智能生成的传闻。使用一个约60000篇Reddit帖子的语料库,分为“可能是自闭症患者”和“普通Reddit”子语料库,比较OpenAI GPT - 2检测模型输出的概率分布。观察子语料库之间的文本特征差异并与人工智能生成文本的特征进行比较。结果显示,两个子语料库中均不到2%的文本被模型标记为人工智能生成,但来自“可能是自闭症患者”子语料库的文本被标记的显著更多。自闭症作者文本特征与人工智能生成文本之间的联系并不直接。人工智能检测模型的广泛使用及其输出中对自闭症作家的潜在偏见引发了伦理审查,作者建议对模型本身及其在学术背景中的使用进行进一步严格审查。
英文摘要
Recent findings suggest that detection models for artificial intelligence (AI) cannot accurately identify AI-generated text and may exhibit bias against certain minority groups. In the present study, anecdotal claims that autistic writers more often have their work flagged as AI-generated are examined empirically. A corpus of approximately 60,000 Reddit posts split into "likely-autistic" and "general-Reddit" subcorpora is used to compare the distribution of probabilities output by the OpenAI GPT-2 detection model. Differences in textual features between subcorpora are observed and compared to reported features of AI-generated text. Results showed that while less than two-percent of either subcorpus was flagged as AI-generated by the model, significantly more texts from the likely-autistic subcorpus were flagged. Connections between features of text with likely-autistic authors and AI-generated text were not straightforward. The widespread use of AI-detection models with a potential bias against autistic writers in their output prompts ethical scrutiny, and the authors recommend further critical examination of the models themselves as well as their use in academic contexts.
CommentsAuthor Accepted Manuscript, Artificial Intelligence in Education 2025
Journal refLecture Notes in Computer Science, 15879, 89-103, Springer Nature Switzerland (2025)
DOI:10.1007/978-3-031-98420-4_7