发表机构
Drexel University; UCLA(德雷塞尔大学; 加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在刻画Reddit上癌症错误信息,引入多维分类法涵盖七个维度。利用专家标注数据评估大语言模型,分析错误信息。发现其约占癌症讨论6%,少样本提示可提高性能,还识别出常见错误信息叙述,为相关建模奠定基础。
AI 中文摘要
社交媒体上与癌症相关的讨论为信息交流和同伴支持提供了重要空间,但也助长了可能影响预防、筛查和治疗决策的错误信息传播。现有关于癌症错误信息的研究往往依赖狭义定义、小规模数据集或二元标签框架。我们引入一种多维分类法,用于刻画Reddit上关于乳腺癌、肺癌、结肠癌和前列腺癌讨论中的癌症错误信息。该分类法涵盖七个维度,包括错误信息存在情况、信息类型、风险水平、立场和主题焦点。利用专家标注数据,我们评估多个大语言模型进行可扩展的错误信息标注,并分析Reddit社区中的癌症错误信息。结果表明,与癌症相关的错误信息约占Reddit上癌症讨论的6%,各社区和错误信息主题存在显著差异。少样本提示显著提高分类性能,特别是对于细微的分类维度。我们还识别出以未经证实的治疗、对传统医学的不信任以及关于诊断和筛查的误导性说法为中心的反复出现的错误信息叙述。我们的分类法、数据集和发现为在线癌症错误信息的多维建模奠定了基础。
英文摘要
Cancer-related discussions on social media provide important spaces for information exchange and peer support, but can also expose users to misinformation with implications for prevention, screening, and treatment decisions. Existing work often treats cancer misinformation as a binary phenomenon, providing limited insight into how misinformation is expressed, engaged with, and associated with potential harm. We introduce a multi-dimensional taxonomy for characterizing cancer misinformation in Reddit discussions of breast, lung, colon, and prostate cancer. Developed through expert annotation, the taxonomy captures seven dimensions spanning misinformation presence, cancer stage, information-seeking and sharing behavior, misinformation type, risk, stance, and topical focus. We evaluate 21 large language models (LLMs) across zero- and few-shot settings and develop a cross-model agreement strategy for scaling misinformation identification to over 133K posts. Our analysis shows that misinformation is heterogeneous in both form and function: unproven and alternative treatments emerge as a prominent topic, misinformation frequently occurs within exchanges that combine information seeking and sharing, and users engage with questionable claims through both endorsement and uncertainty. We further find that classification difficulty varies substantially across dimensions, with risk assessment posing particular challenges for both human annotators and LLMs. Our taxonomy and empirical findings move beyond binary detection toward a more nuanced characterization of how cancer misinformation is produced, discussed, and encountered in online health communities.
CommentsAccepted to Findings of EMNLP 2026