arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18494cs.CL

规模很重要:捷克语 HTML 文档的基础模型

Size Matters: Foundation Model for Czech HTML documents

Martin Dvořák, Vít Tlustoš, Artyom Voronin, Martin Habrovec, Kateřina Podlesná, Barbora Rišová, Josef Vonášek

首次发表
浏览论文内容

中文总结 AI 辅助

HTML-LM 是一个 1.54 亿参数的紧凑基础模型,通过 HTML 感知训练和 ModernBERT 架构,在捷克语网页文档的分类和回归任务上达到最先进水平,并已投入生产。

中文摘要 AI 辅助

在高流量工业环境中创建通用、高质量的网页文档表示,需要既高效又经济的模型。然而,现有方法往往依赖大型模型,忽视 HTML 中固有的结构信息,或受限于短上下文窗口,从而限制了它们处理真实网页的能力。我们提出了 HTML-LM,一个具有 1.54 亿参数的紧凑型基础模型,通过 HTML 感知训练和基于 ModernBERT 的架构解决了这些限制。该模型在 1 亿个网页文档上使用多种目标进行训练,包括掩码语言建模、词袋预测以及从大型语言模型中进行对比蒸馏。因此,HTML-LM 在捷克互联网领域的分类和回归应用中树立了新的最先进水平,超越了更大的编码器和小型 LLM。该模型已投入生产,每秒处理数千个网页文档,并以 CC BY-NC 4.0 许可向社区发布。此 https URL。

英文摘要

Creating universal, high-quality representations of web documents in high-traffic industrial environments requires models that are both performant and economic. Existing approaches, however, often depend on large models, overlook the structural information inherent in HTML, or are constrained by short context windows, limiting their ability to process real-world web pages. We present HTML-LM, a compact foundation model with 154 million parameters that addresses these limitations through HTML-aware training and a ModernBERT-based architecture. It was trained on 100 million web documents using multiple objectives, including masked language modeling, bag-of-words prediction, and contrastive distillation from large language models. Consequently, HTML-LM sets a new state-of-the-art for classification and regression applications in the Czech Internet domain, surpassing both larger encoders and small-sized LLMs. The model is deployed in production, processing thousands of web documents per second, and released to the community under the CC BY-NC 4.0. https://huggingface.co/Seznam/html-lm.

补充信息

↑