AI 中文总结
本文利用信息性度量和大语言模型开发了名为ClickGuard的浏览器扩展,采用混合机器学习架构,结合Transformer嵌入、语言特征和“诱饵性”分数,基于XGBoost建模,能在用户访问前后预警标题党文章,给出可能性分数并解释,还提供文章剧透。
AI 中文摘要
本文介绍了一种由人工智能驱动的浏览器扩展,用于识别标题党,帮助用户避免被误导性的网络文章欺骗。该应用超越了传统检测方法,采用了一种混合机器学习架构,将基于Transformer的嵌入与语言特征和自定义的“诱饵性”分数相结合。在评估了各种自然语言处理技术(从经典向量器到大型语言模型嵌入)之后,开发了一种基于XGBoost的模型,该模型在开放组合数据集上的F1分数达到了91%。最重要的是,该工具可以在用户访问标题党文章之前和之后发出警告。打开文章后,用户会收到一个百分比分数,表明该文章是标题党的可能性。预测基于分析的指标进行解释,包括在所提出的系统中专门开发的指标。该浏览器扩展还提供了一个标题党剧透——文章的一到两句话总结。演示视频:这个https网址}{这个https网址
英文摘要
This paper presents an AI-driven browser extension that identifies clickbait to help users avoid misleading Internet articles. Moving beyond traditional detection, the application employs a hybrid machine learning architecture that combines transformer-based embeddings with linguistically motivated features and a custom "baitness" score. After evaluating various natural language processing techniques -- from classic vectorizers to large language model (LLM) embeddings -- an XGBoost-based model was developed that achieves an F1-score of 91% on the open combined dataset. Most importantly, the tool can warn users before and after they access a clickbait article. After opening an article, the user receives a percentage score indicating the likelihood that it is clickbait. The prediction is explained based on the analyzed metrics, including those specifically developed within the proposed system. The browser extension also provides a clickbait spoiler -- a one- to two-sentence summary of the entire article. Demo video:https://www.youtube.com/watch?v=IJ1gkQV82C4}{https://www.youtube.com/watch?v=IJ1gkQV82C4