面向机器学习应用的网页视觉感知表示
Visual-Aware Representation of Web Pages for Machine Learning Applications
浏览论文内容
中文总结 AI 辅助
本文提出基于FitLayout的网页视觉感知表示与机器学习平台,支持数据集准备、机器学习输入获取,可用于训练图神经网络识别网页关键内容元素,保障结果可复现。
中文摘要 AI 辅助
将机器学习应用于网页具有挑战性,因为需要解释HTML及相关资源,并进行渲染以获得有意义的视觉和布局感知表示。因此,针对网页内容的机器学习探索仍相对不足。本文提出了一个基于开源渲染工具FitLayout的网页视觉感知表示与机器学习平台。该平台提供能够渲染网页的服务器,以基于RDF的表示形式显式捕获网页的视觉和结构属性,并将渲染后的文档持久化存储在集成存储中。处理流程通过REST API控制,同时使用SPARQL查询检索适合作为机器学习算法输入的结构化数据。通过显式建模包含细粒度布局细节的渲染网页,该平台支持数据集共享并保障实验结果的可复现性。其架构支持完整的数据集准备工作流,从网页收集、渲染,到内容元素的预处理与标注,再到下游学习任务。我们还提供了一个Python客户端库,将该平台与标准机器学习工作流集成。作为演示,我们展示了如何将渲染后的网页转换为基于图的表示,并用于训练图神经网络以识别关键内容元素,既说明了该方法的适用性,也证明了结果的可复现性。
英文摘要
Applying machine learning to web pages is challenging due to the need to interpret HTML together with associated resources and perform rendering to obtain a meaningful visual and layout-aware representation. As a result, machine learning over web content remains comparatively underexplored. In this paper, we present a platform for visual-aware representation and machine learning over web pages based on the open-source rendering tool FitLayout. The platform provides a server capable of rendering web pages, explicitly capturing their visual and structural properties in an RDF-based representation, and persisting the rendered documents in an integrated storage. The processing pipeline is controlled via a REST API, while SPARQL queries are used to retrieve structured data suitable as input for machine learning algorithms. By explicitly modeling rendered web pages, including fine-grained layout details, the platform enables dataset sharing and supports the reproducibility of experimental results. The architecture supports the complete dataset preparation workflow, from web page collection and rendering through preprocessing and annotation of content elements to downstream learning tasks. We further provide a Python client library that integrates the platform with standard machine learning workflows. As a demonstration, we show how rendered web pages can be transformed into graph-based representations and used to train graph neural networks for recognizing key content elements, illustrating both the applicability of the approach and the reproducibility of the results.
发表机构
- Brno University of Technology(布尔诺理工大学)
- Faculty of Information Technology(信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。