发表机构
Guizhou University of Finance and Economics; Harbin Institute of Technology (Shenzhen); Guizhou University of Commerce(贵州财经大学; 哈尔滨工业大学(深圳); 贵州商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文从信息融合视角综述了2018年至2024年初的211项虚假评论检测研究,梳理了从传统方法到基于预训练语言模型和大语言模型的技术演进,并指出了对抗生成、跨域迁移等开放挑战。
AI 中文摘要
在线评论塑造了消费者的决策、平台治理和企业的声誉。然而,虚假评论通过向评分系统、推荐流程和公众信任中注入欺骗性证据,破坏了这一信息渠道。大语言模型(LLMs)的兴起从两个方面改变了这一问题。一方面,LLMs能够生成流畅且具有上下文感知能力的欺骗性评论;另一方面,预训练语言模型(PLMs)和LLMs也为检测提供了更强的语义表示。本综述从信息融合的视角回顾了虚假评论检测,涵盖了2018年至2024年初发表的211项研究。我们按证据来源和融合层级对现有工作进行组织,涵盖评论文本、情感、评分行为、时间元数据、用户-产品图、多模态内容、外部知识以及LLM生成的评论。我们追溯了从传统机器学习和深度学习到基于PLM和基于LLM的方法的发展历程,并考察了不同方法如何结合文本、行为、结构和多模态信息。我们还分析了在广泛使用的Amazon、Yelp和OpSpam基准家族上报告的性能趋势,同时指出了由不同标签构建流程、数据划分和评估协议所造成的局限性。最后,我们指出了对抗性生成、跨域迁移、不确定性感知融合、缺失源鲁棒性、可解释性以及针对AI生成欺骗内容的可信评估等方面的开放问题。
英文摘要
Online reviews shape consumer decisions, platform governance, and corporate reputation.Fake reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust mechanisms.The rise of large language models, or LLMs, has changed the problem in two directions.LLMs can generate fluent and context-aware deceptive reviews, while pre-trained language models, or PLMs, and LLMs also provide stronger semantic representations for detection.This survey reviews fake review detection from an information fusion perspective, covering 211 studies published from 2018 to early 2026.We organize existing work by evidence source and fusion level, covering review text, sentiment, rating behavior, temporal metadata, user-product graphs, multimodal content, external knowledge, and LLM-generated signals.We trace the development from traditional machine learning and deep learning to PLM-based and LLM-based methods, and examine how different approaches combine textual, behavioral, structural, and multimodal evidence.We also analyze reported performance trends on widely used Amazon, Yelp, and OpSpam benchmark families, while noting the limitations caused by different label construction procedures, data splits, and evaluation protocols.Finally, we identify open problems in adversarial generation, cross-domain transfer, uncertainty-aware fusion, missing-source robustness, interpretability, and trustworthy evaluation for AI-generated deceptive content.
CommentsFanji Yang and Huiyao Chen contributed equally to this work. Accepted for publication in Information Fusion
DOI:10.1016/j.inffus.2026.104715