发表机构
The University of Edinburgh; Google(爱丁堡大学; 谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种基于LLM的社交媒体分析流水线,从570万条Reddit帖子中检测并分类新兴AI危害,构建了含12类47子类的分类体系,以补充自上而下的AI治理框架。
AI 中文摘要
AI系统的快速部署已造成社会技术、心理和操作层面的危害,这些危害可能逃避事前威胁建模和事后事件追踪。我们引入了一个基于LLM辅助的主题分析流水线,用于从大规模社交媒体数据中动态检测、分类和追踪新兴AI危害。将其应用于2025年1月至2026年6月期间18个月的570万条Reddit帖子摘要,我们整理并发布了一个包含575,000条AI危害相关帖子的数据集,以及一个包含12个类别和47个子节点的自下而上的AI危害分类体系。该分类体系可靠地覆盖了已确立的专家定义风险,同时揭示了自上而下框架所忽视的细粒度危害,例如不同形式的AI隐私侵犯。时间分析揭示了不断演变的以用户为中心的危害,如智能体隐私和安全漏洞、工作场所过早采用AI,以及因AI伴侣服务终止而产生的悲伤。我们的流水线缩短了危害检测时间,并由此补充了朝着更具参与性和响应性的AI治理所做的努力。
英文摘要
The rapid deployment of AI systems has created socio-technical, psychological, and operational harms that can elude ex-ante threat modelling and ex-post incident tracking. We introduce an LLM-assisted thematic analysis pipeline to dynamically detect, categorise, and track emerging AI harms from large-scale social media data. Applying it to 5.7 million Reddit post summaries over 18 months (01/2025 to 06/2026), we curate and release a dataset of 575,000 AI harm-related posts and a bottom-up AI harm taxonomy of 12 categories and 47 subnodes. The taxonomy reliably covers established expert-defined risks while surfacing granular harms that top-down frameworks overlook, such as distinct forms of AI privacy violations. Temporal analysis surfaces evolving user-centric harms, such as agentic privacy and security breaches, premature AI adoption in the workplace, and grief from AI companion discontinuation. Our pipeline shortens harm-detection timelines and hereby complements efforts towards more participatory and responsive AI governance.