BERTopic-病毒传播优先级框架:用于Twitter上COVID-19与猴痘错误信息的主题及对比分析的可扩展框架
BERTopic-Virality Prioritisation: A Scalable Framework for Thematic and Comparative Analysis of COVID-19 and Monkeypox Misinformation on Twitter
- Manchester Metropolitan University(曼彻斯特城市大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出BERTopic-VP框架,结合BERTopic与病毒传播优先级层,搭配混合错误信息检测模块,可高效分析Twitter上COVID-19与猴痘错误信息,能识别高影响力簇,支持监测预警。
AI中文摘要:
大流行期间传播的健康错误信息可迅速获得关注度,形成有害叙事,与公共卫生指导相竞争。大多数主题建模流程将用户互动视为外部结果,限制了其优先选择语义连贯且快速传播主题的能力。我们提出BERTopic-VP,这是一种病毒传播优先级主题建模框架,将基于上下文嵌入的聚类(BERTopic)与事后病毒传播优先级(VP)层相结合。该流程辅以两阶段混合错误信息检测模块,融合基于内容的监督分类器与来自公共卫生知识库的外部验证信号。将该框架应用于COVID-19_FNIR、Monkeypox和Constraint三个基准数据集,其分类性能表现优异,F1值最高达0.950,ROC-AUC最高达0.989,同时在VP阈值为前1%、5%和10%时识别出高影响力簇。对于缺乏原生互动元数据的数据集,优先级基于逻辑传播倾向得分,该得分用作传播潜力的序数代理,而非互动的直接衡量指标。结果表明,整合语义结构、病毒传播感知排名及情感语言分析,可实现对不同大流行错误信息的可扩展且可解释的对比分析。该框架通过呈现低体量但高风险的叙事供分析师审查,支持面向监测的早期预警。
英文摘要:
Health misinformation circulating during pandemics can gain traction rapidly, creating harmful narratives that compete with public health guidance. Most topic-modelling pipelines treat engagement as an external outcome, limiting their ability to prioritise semantically coherent topics that are also rapidly diffusing. We introduce BERTopic-VP, a virality-prioritised topic-modelling framework that combines contextual embedding-based clustering (BERTopic) with a post hoc Virality Prioritisation (VP) layer. The pipeline is complemented by a two-stage hybrid misinformation detection module that fuses a supervised content-based classifier with an external verification signal derived from public-health knowledge bases. Applied to three benchmark datasets, COVID-19_FNIR, Monkeypox, and Constraint, the framework achieves strong classification performance, with F1 up to 0.950 and ROC-AUC up to 0.989, while identifying high-impact clusters under top 1%, 5%, and 10% VP thresholds. For datasets without native engagement metadata, prioritisation is based on a logistic propensity-to-spread score, used as an ordinal proxy for diffusion potential rather than a direct measure of engagement. The results show that integrating semantic structure, virality-aware ranking, and affective-linguistic profiling enables scalable and interpretable comparative analysis of misinformation across pandemics. The proposed framework supports monitoring-oriented early warning by surfacing low-volume but high-risk narratives for analyst review.