面向智能体文档分析的鲁棒层次结构
Robust Hierarchical Structures for Agentic Document Analysis
浏览论文内容
中文总结 AI 辅助
针对LLM智能体忽略文档层次结构的问题,提出SHED两阶段工作流,鲁棒且紧凑地提取结构,提升F-1分数13%-68%,使智能体准确率提高3%-23%且成本降低10倍。
中文摘要 AI 辅助
大型语言模型(LLMs)使我们能够更好地理解文本文档,包括PDF和Word文档。然而,LLMs以及更现代的LLM智能体(即具有工具调用能力的智能体)通常将此类文档视为纯文本,忽略了它们通常按章节和子章节进行层次组织的事实。提取这种结构虽然困难,但可以提高智能体(和人类)的效率和有效性——因为只需处理与给定任务相关的章节。遗憾的是,先前关于结构提取的工作对推断结构与真实结构的匹配程度没有提供形式化保证。相反,我们针对一种鲁棒且紧凑的变体,该变体易于推断且在实践中实用。鲁棒性确保每个子章节标题下的文本是真实结构中相同标题下文本的超集。紧凑性旨在最小化这个超集,从而降低智能体成本(或人类认知负担)。我们提出了SHED,一个用于推断鲁棒且紧凑结构的两阶段工作流。第一阶段是可插拔的,支持无限系列的方法,每种方法都针对特定文档类别保证鲁棒性。我们从理论上使用这些类别及其层次关系来刻画文档空间。实验上,SHED在F-1分数(衡量鲁棒性-紧凑性权衡)上比非LLM基线提高了13%–68%,比昂贵的基于LLM的方法提高了9%–15%。最后,我们展示了SHED推断的结构对智能体文档分析的价值:使用SHED的智能体优于基线,准确率提高3%–23%,同时成本降低高达10倍。
英文摘要
Large Language Models (LLMs) enable us to better understand text documents, including PDFs and Word documents. However, LLMs, as well as more modern LLM agents, i.e., those with tool-calling abilities, typically treat such documents as plain text, ignoring the fact that they are often organized hierarchically into sections and subsections. Extracting this structure, while difficult, can improve efficiency and effectiveness for agents (and humans)---since only sections relevant to a given task need to be processed. Unfortunately, prior work on structure extraction provides no formal guarantees on how well the inferred structure matches the true one. Instead, we target a robust and compact variant that is feasible to infer and useful in practice. Robustness ensures that the text under each subsection header is a superset of the text under the same header in the true structure. Compactness seeks to minimize this superset, reducing agentic cost (or human cognitive load). We propose SHED, a two-stage workflow for inferring a robust and compact structure. The first stage is pluggable with an infinite family of approaches, each guaranteeing robustness for a specific document class. We theoretically characterize the document space using these classes and their hierarchical relationships. Empirically, SHED improves F-1 scores (measuring the robustness--compactness trade-off) by 13%--68% over non-LLM baselines and 9%--15% over expensive LLM-based approaches. Finally, we show how SHED-inferred structures are valuable for agentic document analysis: agents using SHED outperform baselines, achieving 3%--23% higher accuracy while being up to 10x cheaper.
发表机构
- UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。