arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PACE:基于智能体自动化的发布者自适应内容提取

PACE: Publisher-Adaptive Content Extraction via Agentic Automation

Zhanlin Liu, Munirathnam Srikanth

arXiv 2608.27466首次发表:更新:

AI 中文总结

PACE是一种智能体框架,可从代表性页面和用户需求学习发布者特定提取配置,在多模态提取任务中优于可扩展非手动基准,接近手动解析器质量,实现高效自适应内容提取。

AI 中文摘要

Web内容提取对于可靠的大语言模型(LLM)数据管道至关重要,但现有方法往往难以同时满足准确性、可扩展性和适应性的要求。通用提取器可广泛应用,但在发布者特定布局及元数据、图像、表格等更丰富的提取目标上表现脆弱;直接基于LLM的提取灵活性更高,但大规模应用时会产生大量成本和延迟;手动构建的发布者特定解析器可实现高准确率,但需要大量人力进行构建和维护。我们提出PACE,这是一种智能体框架,用于从代表性页面和用户需求中学习发布者特定的提取配置。训练过程中,PACE使用LLM分析页面结构并聚合可复用的提取模式;推理阶段,学习到的配置实例化固定的确定性提取器模板,无需额外LLM调用即可实现可扩展的提取。涵盖文章正文、元数据和多模态提取的实验表明,PACE优于可扩展的非手动基准方法,同时接近手动构建的发布者特定解析器的质量。PACE在文章文本、元数据、图像和表格的提取上表现更强,证明智能体配置学习可实现发布者特定的提取自动化,生成适用于LLM的页面表示,且不仅限于文章文本。

英文摘要

Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and adaptability. General-purpose extractors can be applied broadly, but they are often brittle on publisher-specific layouts and richer extraction targets such as metadata, images, and tables. Direct LLM-based extraction offers greater flexibility, but incurs substantial cost and latency at scale, while manually engineered publisher-specific parsers can achieve high accuracy but require substantial human effort to build and maintain. We introduce PACE, an agentic framework for learning publisher-specific extraction configurations from representative pages and user requirements. During training, PACE uses LLMs to analyze page structure and aggregate reusable extraction patterns. At inference time, the learned configurations instantiate a fixed deterministic extractor template, enabling scalable extraction without additional LLM calls. Experiments spanning article-body, metadata, and multimodal extraction show that PACE outperforms scalable non-manual baselines while approaching the quality of manually engineered publisher-specific parsers. PACE achieves stronger extraction of article text, metadata, images, and tables, demonstrating that agentic configuration learning can automate publisher-specific extraction for LLM-ready page representations beyond article text.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑