发表机构
KAIST; Soongsil University(韩国科学技术院; 崇实大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用Stack Overflow超两百万个问题,以ChatGPT-3.5发布为自然冲击,发现生成式AI导致简单问题骤减、困难问题增多,且简单问题下降集中在数据丰富领域,影响不均衡,侵蚀易获取知识而保留复杂知识。
AI 中文摘要
生成式人工智能(Gen AI)正在重塑个人学习和工作的方式,但其对集体知识——即在线社区共同生产的共享知识体系——的影响仍知之甚少。先前的研究记录了知识共享平台参与度的总体下降,但仍不清楚哪些特定类型的知识最先流失。我们利用Stack Overflow(最大的软件工程在线社区之一)研究这一问题,将ChatGPT-3.5的发布视为自然冲击。通过分析2020年至2025年间发布的超过两百万个问题,我们追踪了集体知识的两个维度——难度和数据可用性——在生成式人工智能发布后的变化。运用多种方法和稳健性检验,我们发现了一致的模式:简单问题急剧减少,而困难问题变得更加普遍,这一模式与代码复杂性的上升相印证。数据丰富的主题和标签失去问题份额,而数据稀缺的主题和标签则获得份额。这两个维度也存在交互作用:简单问题的下降主要集中在数据丰富的领域,而困难问题的增加则与数据可用性无关。这一模式超越了Python,扩展到各种编程语言,且更流行的语言表现出更剧烈的变化。总之,我们的发现揭示了生成式人工智能对集体知识的影响是不均衡的,它首先侵蚀简单、易获取的知识,而更复杂、更不常见的知识则得以存续。
英文摘要
Generative AI (Gen AI) is reshaping how individuals learn and work, but its consequences for collective knowledge, the shared body of knowledge that online communities produce together, remain poorly understood. Prior work has documented an aggregate decline in participation on knowledge-sharing platforms, but it remains unclear which specific kinds of knowledge are being lost first. We study this question using Stack Overflow, one of the largest online communities for software engineering, treating the release of ChatGPT-3.5 as a natural shock. Analyzing over two million questions posted between 2020 and 2025, we track how two dimensions of collective knowledge, difficulty and data availability, change following Gen AI's release. Using diverse methods and robust checks, we find consistent patterns. Easy questions decline sharply while difficult questions become more common, a pattern corroborated by rising code complexity. Data-rich topics and tags lose share of questions, while data-scarce ones gain ground. The two dimensions also interact: the decline in easy questions is concentrated specifically within data-rich domains, while difficult questions increase regardless of data availability. This pattern extends beyond Python across programming languages, with more prevalent languages showing sharper shifts. Together, our findings reveal that Gen AI's impact on collective knowledge is uneven, eroding easy, accessible knowledge first while more complex, less common knowledge persists.