arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

网页价格提取:最新进展与自适应无浏览器实现

Web Price Extraction: State of the Art and an Adaptive Browserless Implementation

Evgeniia Kositsyna, Jorge Lloret-Gazo

arXiv 2609.01030首次发表:更新:

发表机构

University of Zaragoza(萨拉戈萨大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对网页价格提取任务,提出一种自适应无浏览器价格提取系统,通过结合HTML片段化、多维度规则及贝叶斯、遗传算法优化,提升了准确率并降低了处理时间,可作为低计算成本的有竞争力方案。

AI 中文摘要

从网站中提取价格是电子商务市场监测、价格比较和商业分析的关键任务。现有方法大致可分为四类,理解它们在准确性和可扩展性方面的权衡对于选择合适的提取策略至关重要。经典方法依赖手动编写的包装器和从标记页面进行的规则归纳,准确性高但对结构变化的适应性差,且需要大量维护工作。基于浏览器的方法使用Selenium和Puppeteer等工具处理动态JavaScript内容,但消耗大量计算资源且可扩展性差。无浏览器方法通过HTTP请求直接获取HTML,在速度和成本上有显著提升,但依赖针对特定网站校准的规则。基于机器学习和大语言模型(LLM)的方法具有适应性,但需要训练数据和大量计算。本文的主要贡献是一种自适应无浏览器价格提取系统,可提升对网站间结构差异的鲁棒性。我们实现了结合HTML页面片段化与句法、语义及频率规则的基线架构,并通过两种方式对其进行扩展:一种是动态更新规则权重的贝叶斯方法,另一种是优化系统全局参数的遗传算法。该混合方案相较于基线,将准确率从77.2%提升至87.3%,并将单页平均处理时间减少约14%,证实其可作为手动调优的无浏览器解决方案以及资源密集型的基于浏览器或LLM的方法的有竞争力替代方案,能以低计算成本实现高提取准确率。

英文摘要

Price extraction from websites is a key task for market monitoring, price comparison, and business analytics in e-commerce. Existing approaches can be broadly divided into four groups, and understanding their trade-offs in accuracy and scalability is essential for selecting suitable extraction strategies. Classical methods rely on manually written wrappers and rule induction from labeled pages, offering high accuracy but adapting poorly to structural changes and requiring considerable maintenance effort. Browser-based methods, using tools such as Selenium and Puppeteer, handle dynamic JavaScript content but consume large computational resources and scale poorly. Browserless approaches retrieve HTML directly via HTTP requests, offering significant gains in speed and cost, but rely on rules calibrated for specific sites. Methods based on machine learning and large language models offer adaptability but require training data and substantial computation. Our main contribution is an adaptive browserless price extraction system that improves robustness to structural differences between websites. We implemented a baseline architecture combining HTML page fragmentation with syntactic, semantic, and frequency rules, and extended it in two ways: a Bayesian approach that dynamically updates rule weights, and a genetic algorithm that optimizes the system's global parameters. This hybrid scheme increased precision from 77.2% to 87.3% and reduced average per-page processing time by approximately 14% relative to the baseline, confirming it as a competitive alternative to manually tuned browserless solutions and to more resource-intensive browser- or LLM-based methods, offering high extraction accuracy at low computational cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑