发表机构
Kennesaw State University(肯尼索州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究通过两个二元分类任务,用词汇、可读性等特征及多种机器学习模型,区分真实新闻与人类及人工智能生成的假新闻,发现检测人工智能生成的假新闻准确率高,而人类撰写的较难区分,揭示了错误信息检测的不对称性。
AI 中文摘要
随着大语言模型的快速发展,人工智能生成的假新闻与传统人类撰写的错误信息一同出现,引发了关于可检测性是否取决于欺骗性内容来源的问题。本研究通过两个受控二元分类任务来探讨该问题:区分真实新闻与人类撰写的假新闻以及人工智能生成的假新闻。每篇文章用与词汇多样性、可读性和情感特征相关的特征表示,并通过包括逻辑回归、随机森林、支持向量机、梯度提升、神经网络和集成方法在内的多种机器学习模型进行评估。使用接收器操作特征曲线下的面积(AUC)来衡量性能。在所有模型中,人工智能生成的假新闻检测准确率近乎完美,而人类撰写的假新闻与真实新闻更难区分。由于两个任务使用相同的建模流程,这种性能差距反映了文本中的内在统计差异而非方法差异。特征级分析表明,人工智能生成的假新闻表现出更均匀的可读性和情感模式,与真实新闻的重叠较少。这些发现揭示了错误信息检测中的一个关键不对称性:当前方法在识别人工智能生成的内容方面可能非常有效,但对复杂的人类撰写的错误信息仍然不太可靠。因此,检测系统应考虑错误信息的来源,并随着生成模型的发展不断调整。
英文摘要
The rapid advancement of large language models has introduced AI-generated fake news alongside traditional human-written misinformation, raising questions about whether detectability depends on the source of deceptive content. This study examines that issue through two controlled binary classification tasks: distinguishing real news from human-written fake news and from AI-generated fake news. Each article is represented using features related to lexical diversity, readability, and emotional characteristics, and evaluated with several machine learning models, including logistic regression, random forests, support vector machines, gradient boosting, neural networks, and ensemble methods. Performance is measured using the area under the receiver operating characteristic curve (AUC). Across all models, AI-generated fake news is detected with near-perfect accuracy, while human-written fake news is substantially more difficult to distinguish from real news. Because both tasks use the same modeling pipeline, this performance gap reflects intrinsic statistical differences in the text rather than methodological variation. Feature-level analysis shows that AI-generated fake news exhibits more uniform readability and emotional patterns, producing less overlap with real news. These findings reveal a key asymmetry in misinformation detection: current methods may be highly effective at identifying AI-generated content but remain less reliable against sophisticated human-authored misinformation. Detection systems should therefore account for the source of misinformation and continue adapting as generative models evolve.
CommentsAccepted at the International Conference on Machine Learning and Applications (ICMLA 2026)