arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21284cs.CL

新闻爬虫语言模型:一种用于高质量新闻爬虫的小型长上下文模型

news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling

  • Humboldt-Universität zu Berlin(柏林洪堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Pascal Stolzenburg, Jonas Golde, Max Dallabetta, Alan Akbik

AI总结:

研究旨在解决新闻页面结构化内容提取难题,提出基于眼底新闻爬虫库微调的小型长上下文语言模型news-crawler-LM,能将HTML转换为文本和JSON,在HTML到Markdown和JSON提取任务中性能出色,且向研究社区发布了所有模型和工件。

AI中文摘要:

从新闻页面提取结构化内容具有挑战性,因为HTML布局异构、标记不一致且存在大量模板。基于规则的新闻爬虫可通过编码特定网站结构实现高提取精度,但需手动配置以推广到新发布者。大语言模型提供了更灵活的替代方案,但计算成本高限制了实际部署。本文介绍了news-crawler-LM,它是一个小型长上下文语言模型,在眼底新闻爬虫库的高质量、人工验证提取上进行了微调。该模型将原始HTML转换为纯文本和结构化JSON。实验中,在HTML到Markdown和HTML到JSON提取任务中,news-crawler-LM优于强基线,在HTML到纯文本任务中与其他基于规则的解析库相比只有轻微优势。最后将所有模型和工件发布给研究社区。

英文摘要:

Extracting structured content from news pages remains challenging due to heterogeneous HTML layouts, inconsistent markup, and substantial boilerplate such as navigation elements and advertisements. Rule-based news crawlers can achieve high extraction accuracy by encoding site-specific structure, but require manual configuration in order to generalize to new publishers. Large language models provide a more flexible alternative by reducing the need for handcrafted rules, but their high computational cost limits practical deployment. In this paper, we introduce news-crawler-LM, a small long-context language model fine-tuned on high-quality, human-validated extractions from the Fundus news-crawling library. Our model converts raw HTML into plaintext and structured JSON, including fields such as headline, author, publication date, and article body. In our experiments, news-crawler-LM outperforms strong baselines in HTML-to-Markdown and HTML-to-JSON extraction, improving performance by +4.8 BLEU and +6.1 METEOR in the HTML-to-Markdown task, and by +2.2 BLEU and +4.1 METEOR in the HTML-to-JSON task. However, we also observe that our model only slightly better compared to other rule-based parsing libraries on the HTML-to-plaintext task in evaluations on previously unseen publishers. We release all models and artifacts to the research community.

补充信息

↑