AI 中文总结
研究针对现有智能文档处理任务多为独立建模的问题,提出统一智能体系统DocClaw,通过智能体与文档的共享交互处理多种IDP任务,在基准测试中取得有竞争力的性能。
AI 中文摘要
智能文档处理(IDP)涵盖多种任务,包括光学字符识别(OCR)、文档问答(DocQA)和关键信息抽取(KIE)。尽管这些任务目标不同,但它们都需要感知文档内容、获取任务相关信息并逐步优化中间结果。然而,这些任务通常被建模为独立的预测问题,由特定任务的模型或处理流水线解决。我们提出DocClaw,这是一个统一的智能体系统,将多种智能文档处理任务建模为智能体与文档之间的共享交互过程。给定一份文档和特定任务的查询,DocClaw会遵循合适的文档技能,迭代识别所需信息、调用相关工具并将得到的观察结果整合为所需输出。在此过程中,结构化的文档状态会组织可重用的文档知识和特定任务的交互上下文,使智能体能在交互过程中积累、回顾并逐步优化信息。在该建模方式下,特定任务的需求由智能体对查询目标的解读和对应的文档技能捕获,而底层的交互循环、工具空间和文档状态在所有任务间共享。在多个智能文档处理基准上开展的大量实验表明,DocClaw能在单个智能体框架内有效处理多种任务,且与通用视觉语言模型(VLM)及特定任务方法相比,实现了具有竞争力的性能。
英文摘要
Intelligent document processing (IDP) encompasses a broad range of tasks, including optical character recognition (OCR), document question answering (DocQA), and key information extraction (KIE). Despite their distinct objectives, these tasks share a common need to perceive document content, acquire task-relevant information, and progressively refine intermediate results. However, they are typically formulated as separate prediction problems and addressed by task-specific models or processing pipelines. We introduce DocClaw, a unified agentic system that formulates diverse intelligent document processing tasks as a shared process of interaction between an agent and a document. Given a document and a task-specific query, DocClaw follows an appropriate document skill to iteratively identify the information required, invoke relevant tools, and integrate the resulting observations into the desired output. Throughout this process, a structured document state organizes reusable document knowledge and task-specific interaction context, allowing the agent to accumulate, revisit, and progressively refine information as the interaction proceeds. Under this formulation, task-specific requirements are captured by the agent's interpretation of the query objective and the corresponding document skill, while the underlying interaction loop, tool space, and document state are shared across tasks. Extensive experiments across multiple intelligent document processing benchmarks demonstrate that DocClaw effectively handles diverse tasks within a single agentic framework and achieves competitive performance compared with both general-purpose VLMs and task-specific methods.