arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Common Crawl News 与 GDELT 的比较

Comparison of Common Crawl News & GDELT

Ameir El Ouadi, David Beskow

arXiv 2610.00587首次发表:更新:

发表机构

United States Military Academy(美国军事学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文比较了 GDELT 与 Common Crawl News 两个新闻数据集,分析其内容与覆盖范围,发现二者在新闻来源收集地点上存在显著差异。

AI 中文摘要

全球新闻语料对于自然语言处理、知识图谱、大型语言模型以及其他技术工作至关重要。此外,该语料对于理解每天实时互动的人物、地点、组织和事件也具有重要意义。本文比较了当前用于这些任务的两个新闻数据集,即全球事件、语言与语调数据库(GDELT)和 Common Crawl News。我们的研究突出了每个数据集的优势与局限,分析了它们的内容和覆盖范围。值得注意的是,GDELT 依赖来自全球的广播、印刷和网络新闻,而 Common Crawl 则专注于通过网页抓取收集的世界各地新闻网站。我们的分析揭示了这两个数据集在新闻来源收集地点上的显著差异。

英文摘要

The corpus of worldwide news is important for natural language processing, knowledge graphs, large language models, and other technical efforts. Additionally, this corpus is important for understanding the people, places, organizations, and events that interact in real-time every day. This paper compares two news datasets used for these tasks today, namely the Global Database of Events, Language, and Tone (GDELT) and Common Crawl News. Our research highlights the strengths and limitations of each dataset, analyzing their content and coverage. Notably, while GDELT relies on broadcasts, prints, and web news from across the globe, Common Crawl focuses on news sites from around the world gathered through web crawling. Our analysis revealed considerable differences in where the two datasets gather their news sources.

Journal refEl Ouadi, A., & Beskow, D. (2024, April). Comparison of common crawl news & GDELT. In 2024, IEEE International Systems Conference (SysCon) (pp. 1-3). IEEE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑