AI 中文总结
研究自主智能体处理数字文档的能力,引入DocOps评估框架,解构文档操作,评估多种模型,发现其局限性及关键失败模式,揭示能力边界,为健壮无损智能体设计提供启示。
AI 中文摘要
随着自主智能体迅速发展,其可靠处理无处不在的数字文档的能力对于通用人工智能助手和自动化复杂工作流程至关重要。本文引入DocOps,这是一个由分层分类法支撑的可确定性验证的评估框架,该分类法将受现实世界实践启发的文档操作解构为原子维度并提升工作流程复杂性。基于DocOps,我们系统评估了各种智能体框架中的代表性闭源和开源模型,发现即使是最先进的前沿配置在处理高度耦合、远程任务时仍有严重局限性。此外,对现有智能体操作行为的细粒度分析揭示了三个关键失败模式。最终,我们的工作揭示了智能体在维护全局文档一致性方面的能力边界,为复杂数字生态系统中健壮、无损智能体的未来设计提供了启示。
英文摘要
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematically evaluate representative closed- and open-source models across various agentic harnesses, revealing that even the most advanced frontier configurations still exhibit profound limitations when handling highly coupled, long-range tasks. Furthermore, a fine-grained analysis of existing agents' manipulation behaviors uncovers 3 key failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing of structural metadata. Ultimately, our work exposes the capability boundaries of agents in maintaining global document consistency, shedding light on the future design of robust, non-destructive agents for complex digital ecosystems.