AI 中文总结
该研究针对灾害中社交媒体虚假声明严重程度评估问题,提出两阶段框架,结合可信度与危害性维度构建基准,发现LLMs及上下文学习与人工判断对齐度更强。
AI 中文摘要
灾害期间社交媒体上的虚假信息传播迅速,会破坏应急响应工作、公众信任和危机沟通。现有研究主要聚焦于判定社交媒体帖子是否包含虚假信息,但对帖子中嵌入的具体虚假声明以及单个虚假声明的严重程度提供的见解有限。为解决这些局限,我们提出一种评估灾害期间虚假声明严重程度的两阶段框架。第一阶段,我们开发虚假声明提取智能体,从包含文本、图像、视频和链接的多模态社交媒体帖子中识别虚假声明;随后的验证步骤会用支持性证据验证提取出的声明。第二阶段,我们将虚假声明严重程度定义为两个互补维度的结合:可信度(即某一声明被相信的可能性)和危害性(即该声明若被相信可能产生的后果)。我们使用从与飓风和野火相关的Reddit帖子中提取的虚假声明,由人工标注员评估这两个维度,构建声明级严重程度基准。基于该基准,我们将虚假声明严重程度评估作为人机对齐问题展开研究,评估模型是否能在共享评估准则下复现人工判断,而非仅预测严重程度标签。在该基准上的实验显示,传统监督模型与人工判断的对齐度有限,而大型语言模型(Large Language Models, LLMs)的表现显著更强;在所有评估策略中,上下文学习(in-context learning)始终实现与人工判断的最强对齐,凸显了人工示例和共享决策准则对严重程度评估的重要性。
英文摘要
False information spreads rapidly on social media during disasters and can undermine emergency response efforts, public trust, and crisis communication. Existing research primarily focuses on determining whether social media posts contain false information, but provides limited insight into the specific false claims embedded within posts and the severity of individual false claims. To address the limitations, we propose a two-stage framework to assess the severity of false claims during disasters. In the first stage, we develop a false claim extraction agent that identifies false claims from multimodal social media posts containing text, images, videos, and links. A subsequent verification step validates extracted claims with supporting evidence. In the second stage, we define false claim severity as the combination of two complementary dimensions: believability, which determines the likelihood that a claim will be believed, and harmfulness, which captures the potential consequences if it is believed. Human annotators assess both dimensions to construct a claim-level severity benchmark using false claims extracted from Reddit posts related to hurricanes and wildfires. Building upon this benchmark, we investigate false claim severity assessment as a human-AI alignment problem, evaluating whether models can reproduce human judgments under a shared evaluation rubric rather than merely predicting severity labels. Experiments on the benchmark show that traditional supervised models exhibit limited alignment with human judgments, whereas Large Language Models (LLMs) achieve substantially stronger performance. Among the evaluated strategies, in-context learning consistently achieves the strongest alignment with human judgments, highlighting the importance of human examples and shared decision criteria for severity assessment.
CommentsThe 2026 ACM Conference on Human-AI Complementarity and Alignment (HCOMP 26)