我们还能信任灾害社会感知吗?关于检测AI生成社交媒体帖子的实证证据
CrisisFake: Benchmark Validity of AI-Generated Text Detection for Disaster Social Sensing
- The University of Sydney(悉尼大学)
- University of Minnesota Twin Cities(明尼苏达大学双城分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过构建12,000条灾害相关帖子数据集,评估多种文本AI检测器,发现其检测AI生成帖子的可靠性不足,建议依靠多模态和上下文证据保障灾害社会感知的信任。
AI中文摘要:
灾害社会感知将公众社交媒体帖子转化为态势感知和人道主义需求的证据,但生成式人工智能(AI)可以生成看似目击者报告的合理信息。本研究调查基于文本的AI检测器能否可靠地区分人类撰写的和AI生成的灾害帖子。我们构建了一个包含12,000条文本的数据集,组织为来自九次灾害的3,000个匹配语义单元:原始人类帖子(H0)、经LLM最小程度校对的人类帖子(H1)、基于相同已验证事实的事实性AI生成帖子(A0),以及这些AI帖子的情感框架版本(A1)。一个包含来自42个事件的6,000条文本的独立语料库支持模型选择和阈值校准。我们评估了OSM-Det、Fast-DetectGPT、Binoculars和直接大型语言模型(LLM)评判器,涵盖五个模型家族,然后测试了灾害领域校准、冻结编码器线性读出、配对变换敏感性和数据集伪影控制。在十四个冻结跨家族配置中,AUROC为0.402-0.517,在基于校准的低误报率操作点下,最佳前瞻召回率为3.6%;OSM-Det在实现6.7%误报率时达到AUROC 0.521和10.4%召回率。一个灾害训练的线性头达到AUROC 0.817,但一个七特征表面分类器在H0对A0对比中达到0.784,而中和已识别的表面不对称性将线性头从0.733降至0.594。该线性头还能区分A0和A1,尽管来源不变。结果表明,基于文本的检测不足以作为运营信任门;应依靠多模态声明、可问责来源和其他上下文证据来保障灾害社会感知中的信任。
英文摘要:
Disaster social sensing converts public social-media posts into evidence for situational awareness and humanitarian response, but plausible LLM-generated posts can contaminate this information stream and distort assessments of needs, damage, and resource priorities. This study empirically investigates whether text-based detectors can distinguish human-authored from LLM-generated disaster posts and what textual cues underlie their judgments. We construct CrisisFake, a Qwen2.5-7B-based dataset of 12,000 texts organized into 3,000 matched semantic units from nine disasters. Each unit contains an original human post, a minimally LLM-proofread human post, a fact-preserving LLM-generated post produced using LoRA, and an affectively reframed version of the artificial post. A separate 6,000-text corpus spanning 42 disaster events is independently constructed to support model selection and threshold calibration. We evaluate OSM-Det, Fast-DetectGPT, Binoculars, and direct LLM judges across five model families, and further examine how disaster-domain calibration and superficial linguistic cues, such as retweet markers, user mentions, hashtags, URLs, punctuation, and text length, affect detector performance. Across fourteen frozen configurations, AUROC ranges from 0.402 to 0.517, while the best prospective recall at a calibration-derived low-false-positive operating point is only 3.6%, indicating near-chance discrimination. For OSM-Det, a disaster-calibrated linear head improves AUROC to 0.817; however, a seven-feature textual classifier alone reaches AUROC 0.784 on the original-versus-factual-LLM contrast, and neutralizing identified surface asymmetries reduces the corresponding linear-head AUROC from 0.733 to 0.594. These findings provide empirical evidence that LLM-generated text detection is largely driven by linguistic cues and remains insufficiently robust for short-form disaster social media.