AI 中文总结
Scout是一款基于文档相似度的可扩展文档提取工具,通过生成优化规则集实现高性价比数据提取,在多个真实数据集上达到与前沿LLM相当的准确率,成本大幅降低且优于同类基于程序的方法。
AI 中文摘要
从大型文档集合中提取数值是跨多个领域数据分析的核心支撑。前沿大语言模型(LLM)能准确提取这类数值,但用其处理整个集合的成本高得难以承受,而这种成本大多是可以避免的:现实世界的文档集合存在丰富的相似性,因此针对同一查询,相似文档的答案往往出现在相似位置,LLM只需读取该小范围内容,无需读取整个文档。此前利用这种相似性的方法存在不足:要么假设文档结构固定,要么假设答案是输入的一组子串,并使用LLM生成的程序直接返回答案。即使是前沿智能体也无法生成有效程序来直接定位答案的范围,因为搜索空间大,且从小样本学习到的程序往往会过拟合。我们提出Scout,一种能生成准确且高性价比的程序(我们称之为规则)以大规模提取数据的工具。从少量采样文档中,Scout会生成广泛的规则集,并通过选择帕累托最优子集来优化,该子集成本低且不牺牲准确性。我们证明规则优化是NP难问题,并给出具有可证近似保证的贪心解决方案。Scout可处理仅部分相似的集合,其中相似性存在于文档簇内。在这种情况下,一种无需LLM的采样策略会从每个簇中抽取样本,而级联策略会选择优化后的规则子集,当所选规则不包含答案时则回退到未优化的规则集。在六个真实数据集上的实验表明,Scout的准确性与最强基线(读取每个完整文档的前沿LLM智能体)相当,在1000个文档的集合上成本低61倍至超过1000倍,且比最强的基于程序的先前方法准确率高61%。
英文摘要
Extracting values from large document collections powers data analysis across many domains. Frontier LLMs extract such values accurately, but processing an entire collection with one is prohibitively costly. Yet this cost is largely avoidable: real-world collections exhibit rich similarity, so for the same query over similar documents, the answer tends to recur in similar locations; an LLM need only read that small span, not the whole document. Prior methods that exploit this similarity fall short: they either assume a rigid document structure, or assume the answer is a set of substrings of the input and use an LLM-generated program to return it directly. Even a frontier agent fails to generate effective programs to directly locate the answer's span, as the search space is large and programs learned from a small sample tend to overfit. We present Scout, a tool that generates accurate and cost-effective programs (that we call rules) to extract data at scale. From a few sampled documents, Scout generates a broad rule set and refines it by selecting a pareto-optimal subset with low cost without sacrificing accuracy. We prove rule refinement is NP-hard and give a greedy solution with a provable approximation guarantee. Scout handles collections that are only partly similar, where similarity holds within clusters of documents. In this setting, a sampling strategy, using no LLM, draws samples from each cluster; and a cascade strategy selects a subset of refined rules, falling back to the unrefined rule set when the selected rules don't contain the answer. Experiments on six real-world datasets show that Scout matches the accuracy of the strongest baseline, a frontier LLM agent that reads each full document, while being 61x to over 1000x cheaper on a collection of 1,000 documents, and is 61% more accurate than the strongest prior program-based approach.