新闻爬虫语言模型:一种用于高质量新闻爬虫的小型长上下文模型
news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling
- Humboldt-Universität zu Berlin(柏林洪堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究旨在解决新闻页面结构化内容提取难题,提出基于眼底新闻爬虫库微调的小型长上下文语言模型news-crawler-LM,能将HTML转换为文本和JSON,在HTML到Markdown和JSON提取任务中性能出色,且向研究社区发布了所有模型和工件。
AI中文摘要:
从新闻页面提取结构化内容具有挑战性,因为HTML布局异构、标记不一致且存在大量模板。基于规则的新闻爬虫可通过编码特定网站结构实现高提取精度,但需手动配置以推广到新发布者。大语言模型提供了更灵活的替代方案,但计算成本高限制了实际部署。本文介绍了news-crawler-LM,它是一个小型长上下文语言模型,在眼底新闻爬虫库的高质量、人工验证提取上进行了微调。该模型将原始HTML转换为纯文本和结构化JSON。实验中,在HTML到Markdown和HTML到JSON提取任务中,news-crawler-LM优于强基线,在HTML到纯文本任务中与其他基于规则的解析库相比只有轻微优势。最后将所有模型和工件发布给研究社区。
英文摘要:
Extracting structured content from news pages remains challenging due to heterogeneous HTML layouts, inconsistent markup, and substantial boilerplate such as navigation elements and advertisements. Rule-based news crawlers can achieve high extraction accuracy by encoding site-specific structure, but require manual configuration in order to generalize to new publishers. Large language models provide a more flexible alternative by reducing the need for handcrafted rules, but their high computational cost limits practical deployment. In this paper, we introduce news-crawler-LM, a small long-context language model fine-tuned on high-quality, human-validated extractions from the Fundus news-crawling library. Our model converts raw HTML into plaintext and structured JSON, including fields such as headline, author, publication date, and article body. In our experiments, news-crawler-LM outperforms strong baselines in HTML-to-Markdown and HTML-to-JSON extraction, improving performance by +4.8 BLEU and +6.1 METEOR in the HTML-to-Markdown task, and by +2.2 BLEU and +4.1 METEOR in the HTML-to-JSON task. However, we also observe that our model only slightly better compared to other rule-based parsing libraries on the HTML-to-plaintext task in evaluations on previously unseen publishers. We release all models and artifacts to the research community.