arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.31082cs.AIcs.CLcs.DB

通过非结构化数据的自适应结构化实现 token 高效的数据推理智能体

Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出智能体数据拆解方法,通过自适应推测性结构化非结构化数据,降低 LLM 智能体推理成本,在 FanOutQA 基准上实现 53% 的成本削减且保持准确率,为智能体推理数据基础设施提供了初步方案。

中文摘要 AI 辅助

有价值的数据仍嵌入在网页、报告、合同、文件、财报电话会议和 PDF 等非结构化来源中。企业 AI 的重大赌注是部署大语言模型(LLM)智能体,使其能对这些数据进行推理,为每位知识工作者解答复杂问题。如今智能体可以做到这一点,但成本高得令人望而却步:每个问题都会反复打开大型文档以收集分散的证据,消耗多达 100 万个 token。然而,如果数据已经结构化,同一个问题就会简化为廉价的数据库查询。例如,在 FanOutQA 基准测试中,对理想预结构化存储进行推理的成本低 28 倍,且当问题涉及更多文档时,成本差距会扩大至数个数量级。但预先对所有内容进行结构化是不可行的:文档包含的可能结构远多于任何工作负载会用到的结构,且有用的结构和文档要到查询到来时才会知晓。我们提出了智能体数据拆解(agentic data cracking)方法,该方法在推理过程中作为副产品,自适应且推测性地对非结构化数据进行结构化。结构化是自适应的,因为观察到的查询会决定何时进行结构化以及什么内容重要;同时它也是推测性的,因为其范围超出当前问题。每当智能体为解答问题而打开文档时,一个拆解子智能体就会从已加载的上下文以边际成本分叉出来,提取可能服务于未来相关查询的基于事实的结构。随着时间推移,越来越多的查询被结构化数据完全覆盖,无需打开文档即可解答,同时将智能体的准确率保持在接近检索增强生成(RAG)的成本水平。在 FanOutQA 上,仅为每个测试问题增加一个相关问题,拆解方法就降低了 53% 的成本,同时保持了准确率。智能体数据拆解是迈向用于非结构化数据推理的下一代数据基础设施的第一步:它是模型下方的共享底层结构,智能体推理已付出代价所挖掘的知识会在此积累。

英文摘要

Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions for every knowledge worker. Agents can do this today, but at prohibitive cost. Each question repeatedly opens large documents to recover scattered evidence, consuming up to a million tokens. However, if the data were already structured, the same question would reduce to a cheap database lookup. For example, on FanOutQA benchmark, reasoning over an ideal pre-structured store is 28X cheaper, and the gap grows to orders of magnitude as questions fan out over more documents. Yet structuring everything in advance is not viable: documents hold vastly more possible structure than any workload will use, and the useful structure and documents are unknown until queries arrive. We propose agentic data cracking, a method that structures unstructured data adaptively and speculatively as a byproduct of reasoning itself. Structuring is adaptive because observed queries decide when it happens and what matters, and speculative because it goes beyond the current question. Whenever the agent opens a document to answer, a cracking sub-agent forks from the already-loaded context at marginal cost and extracts grounded structure likely to serve related future queries. Over time, an increasing share of queries is fully covered by structured data and answered without opening a document, keeping agentic accuracy at close to RAG cost. On FanOutQA, extended with merely one related question per test question, cracking cuts cost by 53% while preserving accuracy. Agentic data cracking is a first step toward next-generation data infrastructure for agentic reasoning over unstructured data: a shared substrate beneath the model where knowledge that reasoning already paid to uncover accumulates.

发表机构

  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑